TodayAI

资源库

推理与服务 Open Source

面向模型推理、本地运行与高吞吐服务部署的开源引擎。

AWQ

mit-han-lab/llm-awq

激活感知权重量化的开源实现。

quantizationinferencellm
推理与服务PythonMIT

BentoML

bentoml/BentoML

面向模型服务与部署的开源平台。

servingmlopsdeployment
推理与服务PythonApache-2.0

Candle

huggingface/candle

Hugging Face 用 Rust 实现的机器学习推理框架。

inferencerustml
推理与服务RustMIT

Cog

replicate/cog

将机器学习模型打包为标准容器的工具。

servingcontainersml
推理与服务GoApache-2.0

CTranslate2

OpenNMT/CTranslate2

面向 Transformer 模型的快速推理引擎。

inferencetransformersllm
推理与服务C++MIT

FastChat

lm-sys/FastChat

面向训练、服务与评测聊天模型的开源平台。

servingchatevaluation
推理与服务PythonApache-2.0

GGML

ggerganov/ggml

面向张量计算与本地模型运行的张量库。

inferencelocal-airuntime
推理与服务CMIT

GPTQ

IST-DASLab/gptq

面向 LLM 的后训练量化方法实现。

quantizationinferencellm
推理与服务PythonApache-2.0

KoboldCPP

LostRuins/koboldcpp

基于 llama.cpp 的本地推理与 UI 前端。

inferencelocal-aiui
推理与服务C++AGPL-3.0

llama.cpp

ggerganov/llama.cpp

在消费级硬件上高效运行 LLM 的 C/C++ 推理实现。

inferencelocal-aillm
推理与服务C++MIT

LocalAI

mudler/LocalAI

兼容 OpenAI API 的本地模型推理网关。

local-aiinferenceapi
推理与服务GoMIT

MLC LLM

mlc-ai/mlc-llm

面向多端设备的 LLM 编译与部署项目。

inferencelocal-aicompiler
推理与服务PythonApache-2.0

MLX

ml-explore/mlx

Apple 面向 Apple Silicon 的机器学习框架。

inferenceappleml
推理与服务C++MIT

Ollama

ollama/ollama

本地运行与管理大模型的开源工具。

local-aiinferencellm
推理与服务GoMIT

ONNX Runtime

microsoft/onnxruntime

跨平台机器学习推理加速运行时。

inferencemlruntime
推理与服务C++MIT

OpenLLM

bentoml/OpenLLM

用于在生产环境服务开源 LLM 的平台。

inferenceservingllm
推理与服务PythonApache-2.0

SGLang

sgl-project/sglang

面向结构化生成与高吞吐推理的开源运行时。

inferenceservinggeneration
推理与服务PythonApache-2.0

TensorRT-LLM

NVIDIA/TensorRT-LLM

NVIDIA 面向 LLM 优化推理的开源库。

inferencenvidiallm
推理与服务C++Apache-2.0

Text Generation Inference

huggingface/text-generation-inference

Hugging Face 面向文本生成模型的生产推理服务。

inferenceservingllm
推理与服务PythonApache-2.0

vLLM

vllm-project/vllm

面向大语言模型高吞吐推理与服务部署的开源引擎。

inferenceservingllm
推理与服务PythonApache-2.0