lorax

Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

Fork A fork of predibase/lorax; the README may describe the upstream project.

Overview

LoRAX: Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

LoRAX (LoRA eXchange) is a framework that allows users to serve thousands of fine-tuned models on a single GPU, dramatically reducing the cost of serving without compromising on throughput or latency.

Serving a fine-tuned model with LoRAX consists of two components:

LoRAX supports a number of Large Language Models as the base model including Llama (including CodeLlama), Mistral (including Zephyr), and Qwen. See Supported Architectures for a complete list of supported base models.

Base models can be loaded in fp16 or quantized with bitsandbytes, GPT-Q, or AWQ.

Supported adapters include LoRA adapters trained using the PEFT and Ludwig libraries. Any of the linear layers in the model can be adapted via LoRA and loaded in LoRAX.

The minimum system requirements need to run LoRAX include:

See OpenAI Compatible API for details.

From the project’s README on GitHub.

At a glance

RepositoryKentucky-Open-Science/lorax
Research areaForks of other projects
Primary languagePython
LanguagesPython 70.2%, Rust 15.9%, Cuda 10.8%, C++ 1.9%, Dockerfile 0.4%, Shell 0.4%
LicenseApache-2.0
Stars / forks0 / 0
Open issues and pull requests17
Created2024-07-02
Last push2026-04-16
Default branchmain
Homepagehttps://loraexchange.ai
Forked frompredibase/lorax — Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

What the README covers

Top contributors

Get the code

git clone https://github.com/Kentucky-Open-Science/lorax.git