Model Gallery

8 models from 1 repositories

Filter by type:

Filter by tags:

qwen3-4b-vllm-cpp
Qwen3-4B on vllm.cpp, in bf16. The small end of the engine's gated dense family, which reaches parity with vLLM on every axis at concurrency 1. bf16 rather than NVFP4 on purpose: this is the entry that runs where the flagship NVFP4 checkpoints cannot, including Apple Silicon via Metal, Vulkan and plain CPU. Roughly 8 GB of weights, plus about 4.5 GB of KV cache at the context configured here. Tool calling and the thinking split are parsed inside the engine.

Repository: localaiLicense: apache-2.0

qwen3-tts-llamacpp
Qwen3-TTS 1.7B Base served by the llama.cpp backend, using upstream's own GGUF conversion. Runs on the full llama-cpp accelerator matrix (CUDA, ROCm, SYCL, Vulkan, Metal). Streaming output and zero-shot voice cloning: set `voice` to a reference clip or a saved Voice Library profile, which is required since the Base checkpoint has no built-in speaker. 24kHz mono, 10 languages. Q8_0 backbone (~1.8 GB) plus a Q8_0 projector.

Repository: localaiLicense: apache-2.0

deepseek-v4-flash-q2
DeepSeek V4 Flash (IQ2XXS GGUF, ~81 GB) - only loadable via the ds4 backend. Requires >=128 GB RAM. Metal (Darwin) or CUDA (Linux). See https://github.com/antirez/ds4 for details.

Repository: localai

deepseek-v4-flash-q2-q4
DeepSeek V4 Flash (mixed q2/q4 GGUF, ~91 GB) - only loadable via the ds4 backend. The last 6 expert layers are kept at Q4_K (the rest IQ2XXS), trading a little extra memory for higher quality than the pure-q2 build while still fitting in RAM on a 128 GB machine. imatrix-tuned. Metal (Darwin) or CUDA (Linux). See https://github.com/antirez/ds4 for details.

Repository: localai

deepseek-v4-flash-q4-ssd
DeepSeek V4 Flash (full 4-bit experts GGUF, ~153 GB) - only loadable via the ds4 backend, with SSD streaming enabled so it runs on a 128 GB machine even though the weights do not fit in RAM: routed MoE experts stream from the GGUF on SSD while the non-routed weights stay resident. SSD streaming is Metal (Darwin) only; generation speed depends on SSD speed and the expert cache. Tune the routed-expert cache with the 'ssd_streaming_cache_experts:NGB' option (default: automatic budget). See https://github.com/antirez/ds4.

Repository: localai

deepseek-v4-flash-q2-mtp
DeepSeek V4 Flash (IQ2XXS GGUF, ~81 GB) paired with the optional MTP speculative-decoding weights (~3.5 GB) for a slight speedup. Only loadable via the ds4 backend; requires >=128 GB RAM. MTP helps only with greedy decoding (temperature 0), so the override pins temperature to 0. Metal (Darwin) or CUDA (Linux). See https://github.com/antirez/ds4 for details.

Repository: localai

deepseek-v4-pro-q2-ssd
DeepSeek V4 Pro (IQ2XXS GGUF, ~433 GB, imatrix-tuned) - only loadable via the ds4 backend, with SSD streaming so the Pro-class model can be run on a 128 GB machine. This is experimental and slow: it needs ~433 GB of free SSD plus enough RAM for the resident weights, KV cache, and routed-expert cache, and is best used with thinking off for inspection or occasional work. SSD streaming is Metal (Darwin) only. See https://github.com/antirez/ds4.

Repository: localai

nimble-9b-vllm-cpp
Nimble is a decision model from Bespoke Labs (the model Ollama serves as nimble). For each question it runs one forward pass and reads the logits of the answer letters at the last prompt position. It answers typed choice, noul and score questions about a text state and does not generate text. This entry installs a converted redistribution of bespokelabs/Bespoke-Nimble-9B: the LoRA adapter is merged into its Qwen3.5-9B base in BF16 and config.json names the NimbleModel architecture. The checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers and vLLM cannot load it. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp project compared this converted directory on CPU against the Bespoke authors' own code over seven questions: 7 of 7 answers equal, largest probability difference 0.0054. This entry was installed and served through LocalAI on CPU and gave the answers shown in the model card example. There is no accuracy benchmark and no GPU (CUDA, ROCm, Metal) run. The engine refuses fields with more than 26 choices (upstream allows 255). The entry sets an 8192-token context (Nimble's own prompt limit) and a KV pool that holds four such sequences. BF16 weights, about 19.3 GB, pinned to a revision. On CPU the model needs about 20 GB of free RAM (measured peak 18.4 GB resident). On a 20-thread CPU a request with three questions (988 prompt tokens) took 30 seconds warm.

Repository: localaiLicense: apache-2.0