riton-inference-server/backend: Common source, scripts and utilities for creating Triton backends.
github.com · 3,379 words · saved by 1 readers
Common source, scripts and utilities for creating Triton backends.
Triton Inference Server Backend A Triton backend is the implementation that executes a model. A backend can be a wrapper around a deep-learning framework, like PyTorch, TensorFlow, TensorRT or ONNX Runtime. Or a backend can be custom C/C++ logic performing any operation (for example, image pre-processing). This repo contains documentation on Triton backends and also source, scripts and utilities for creating Triton backends. You do not need to use anything provided in this repo to create a Triton backend but you will likely find its contents useful. Frequently Asked Questions Full documentatio
related reading
- server/docs/user_guide/model_configuration.md at main · triton-inference-server/server · GitHubgithub.com
- Together AI | The AI Native Cloudtogether.ai
- tutorials/Conceptual_Guide/Part_1-model_deployment at main · triton-inference-server/tutorials · GitHubgithub.com
- tutorials/Conceptual_Guide/Part_4-inference_acceleration/README.md at main · triton-inference-server/tutorials · GitHubgithub.com
- tutorials/HuggingFace at main · triton-inference-server/tutorials · GitHubgithub.com
- API Reference — TensorRT LLMnvidia.github.io
- Accelerated Inference for Large Transformer Models Using NVIDIA Triton Inference Server | NVIDIA Technical Blogdeveloper.nvidia.com
- Inference Platform: Deploy AI models in production | Basetenbaseten.co
- GitHub - JustVugg/colibri: Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦github.com
- Tinkerthinkingmachines.ai
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com