Model: Make loading use less VRAM

The model loader was using more VRAM on a single GPU compared to base exllamav2's loader. This was because single GPUs were running using the autosplit config which allocates an extra vram buffer for safe loading. Turn this off for single-GPU setups (and turn it off by default). This change should allow users to run models which require the entire card with hopefully faster T/s. For example, Mixtral with 3.75bpw increased from ~30T/s to 50T/s due to the extra vram headroom on Windows. Signed-off-by: kingbri <bdashore3@proton.me>
2026-04-28 02:01:24 +00:00 · 2024-02-06 22:29:56 -05:00
parent fedebadc81
commit 849179df17
3 changed files with 26 additions and 9 deletions
--- a/OAI/types/model.py
+++ b/OAI/types/model.py
@@ -70,7 +70,7 @@ class ModelLoadRequest(BaseModel):
        default=None,
        examples=[4096],
    )
-    gpu_split_auto: Optional[bool] = True
+    gpu_split_auto: Optional[bool] = False
    gpu_split: Optional[List[float]] = Field(
        default_factory=list, examples=[[24.0, 20.0]]
    )