Tufalabs
Pinned Loading
Repositories
- vllm Public Forked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- transformers Public Forked from huggingface/transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
- verl Public Forked from verl-project/verl
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
- transformer-engine Public Forked from NVIDIA/TransformerEngine
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
- sglang Public Forked from sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
- flash-attention Public Forked from Dao-AILab/flash-attention
Fast and memory-efficient exact attention
- tufazip Public
Top languages
Loading…
Most used topics
Loading…
