TiloBox
Back to directory
TensorRT-LLM project preview

TensorRT-LLM

NVIDIA’s open-source library for optimizing and accelerating Large Language Model inference on GPUs.

LicenseApache-2.0
GitHub stars9.5k
Last commit1 weeks ago
Tags6 topics
Nvidia InferenceTensor ParallelismGpu AccelerationArtificial IntelligenceIn Flight BatchingFp8 Quantization
Overview

Why consider TensorRT-LLM?

TensorRT-LLM is NVIDIA’s official open-source inference acceleration library for LLMs on NVIDIA GPUs. Built on TensorRT deep learning compilers, it implements FP8/INT4 quantization, in-flight batching, paged KV caching, and multi-GPU tensor parallelism for massive inference throughput.

Guided learning

Learn TensorRT-LLM by building

Practical setup notes, real use cases, and copy-ready examples in one focused guide.

1 min read 2 sections
In this guide2 sections

Overview of TensorRT-LLM

TensorRT-LLM maximizes token generation speeds on NVIDIA H100, A100, and RTX GPUs with custom CUDA kernels.

Quickstart

bash
1pip install tensorrt_llm -U --pre --extra-index-url https://pypi.nvidia.com

TensorRT-LLM is licensed under the Apache License Version 2.0.

Related tools

More options with a similar category or technology profile.

TensorRT-LLM FAQs

TensorRT-LLM is listed as a Developer Tools tool on TiloBox. Review the overview, features, and official documentation on this page to decide whether it solves your specific workflow.

Start with the project's GitHub repository and official website for supported installation and deployment instructions. Test the setup with representative data or a small project before rolling it out more widely.

TensorRT-LLM is listed under the Apache-2.0 license. Read the complete license text and the project's notices before using, modifying, or distributing the software.

Production readiness depends on your requirements. Review maintenance activity, security practices, documentation, backup and upgrade procedures, and compatibility with your stack; then validate it in a non-production environment.

TensorRT-LLM is listed as an alternative to vLLM. Compare the core workflow, deployment model, integrations, and licensing against your must-have requirements before switching.