Model Gallery

24 models from 1 repositories

Filter by type:

Filter by tags:

qwen3.8-27b-exl3-vllm-cpp
Qwen3.8-27B EXL3 3.5bpw served by vllm.cpp, LocalAI's C++ vLLM-style runtime. The published checkpoint generates on CUDA and measured 16.7 tokens/s on GB10 with the pinned revision and the limits configured here. This is the target-only setup. Use the DFlash2 variant for the measured speculative-decoding configuration. The entry downloads the complete revision-pinned repository, including its configuration, tokenizer, index, and safetensors shards.

Repository: localai

qwen3.8-27b-dflash2-exl3-vllm-cpp
Qwen3.8-27B EXL3 with its EXL3 DFlash2 companion, served by vllm.cpp. On GB10 this pinned pair measured 48.7 tokens/s at a seven-token draft budget, versus 16.7 tokens/s target-only, with token-identical greedy output. LocalAI stages both complete Hugging Face repositories before load. The content-addressed companion snapshot is passed to the backend as the draft model, while the speculative method and seven-token budget remain fixed.

Repository: localai

deepseek-v4-flash-spark-exl3-vllm-cpp
DeepSeek V4 Flash's Spark and GB10-oriented REAP-K216 EXL3 checkpoint, served by vllm.cpp. It needs CUDA and roughly 100 GiB for its large rank-sliced checkpoint. The repository is pinned to its latest recorded revision. vllm.cpp's existing runtime evidence measured the older 22f28d32b9b29b4352eaa380ff8c2c170b2847ab revision; this entry does not claim that the newer revision has passed the same end-to-end gate.

Repository: localai

deepseek-v4-flash-exl3-3bpw-vllm-cpp
Experimental non-Spark DeepSeek V4 Flash EXL3 3.0bpw checkpoint served by vllm.cpp. This is the complete, non-REAP layout and requires a large multi-GPU CUDA system. The publisher describes the artifact as structurally complete but has not passed end-to-end generation. Treat this entry as an integration target, not as a correctness- or performance-gated configuration.

Repository: localai

qwen3.6-27b-nvfp4-vllm-cpp
Qwen3.6-27B in NVFP4, served by vllm.cpp: LocalAI's own C++ port of vLLM, with no Python at inference time. This is the reference text-generation checkpoint the engine is gated on, token-for-token identical to vLLM's own greedy output over the 235-prompt correctness battery, and measured at or above vLLM's throughput at every concurrency from 1 to 32. The weights are PINNED to revision 890bdef7. That pin is load-bearing, not housekeeping: the same repository name was later re-quantized to FP8 W8A8 throughout, so an unpinned copy of this entry serves entirely different weights with no error and none of the measured behaviour above. Needs a Blackwell-class NVIDIA GPU (NVFP4 has no kernel on older architectures) and roughly 25 GB of weights plus KV cache. Tool calling and the thinking split are parsed inside the engine.

Repository: localaiLicense: apache-2.0

qwen3.6-27b-nvfp4-mtp-vllm-cpp
Qwen3.6-27B NVFP4 on vllm.cpp with MTP speculative decoding enabled. MTP (Multi-Token Prediction) drafts from a head that ships inside the target checkpoint's own mtp.* tensors, so there is no second model to download and no extra weights to manage. The verifier accepts roughly 85% of drafted tokens on prose and 92% on code, worth about 1.5x to 1.6x the decode throughput of the same weights with speculation off, and it holds that lead at concurrency 2, 4 and 8. Same weights and same revision pin as qwen3.6-27b-nvfp4-vllm-cpp; install that entry instead if you would rather not spend the extra memory. The speculative state (a doubled recurrent-state slot plus the draft cache and head) costs roughly 3.6 GB on top of the base footprint. MTP here is depth 1 by construction: the engine refuses num_speculative_tokens above 1 for this method.

Repository: localaiLicense: apache-2.0

qwen3.6-27b-nvfp4-dflash-vllm-cpp
Qwen3.6-27B NVFP4 on vllm.cpp with DFlash block-diffusion speculative decoding: the fastest configuration of this model the engine ships. Where MTP drafts one token at a time, DFlash drafts a whole 16-token block in a single non-autoregressive pass from a separate 3.5 GB drafter, then the target verifies the block in one step. At concurrency 1 that measures 2.9x the throughput of the same weights with speculation off, and at or above vLLM's own DFlash-on decode. Both checkpoints are installed for you: the target as a revision-pinned snapshot, the drafter into models/Qwen3.6-27B-DFlash, which is where the backend looks when speculative_config.model names it. The drafter shares the target's embed_tokens and lm_head, so the two are not independently swappable. Needs a Blackwell-class NVIDIA GPU and roughly 28 GB of weights in total.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-nvfp4-vllm-cpp
Qwen3.6-35B-A3B in NVFP4, served by vllm.cpp. A 35B mixture-of-experts model with roughly 3B parameters active per token, so it reads like a much larger model while costing about as much per token as a small one. Image input is implemented in the engine but is not token-gated against vLLM yet, so the vision usecase on this entry is experimental. This is the engine's gated MoE checkpoint: token-for-token identical to vLLM over the 315-prompt battery on both the synchronous and asynchronous paths, at 0.92x to 0.97x vLLM's throughput from concurrency 1 to 32. The architecture is a gated-delta-net hybrid, so automatic prefix caching is off by default here where it would be on for a dense model. That is the engine's own default and this entry does not override it. Needs a Blackwell-class NVIDIA GPU and roughly 23 GB of weights plus KV cache.

Repository: localaiLicense: apache-2.0

qwen3.6-35b-a3b-nvfp4-mtp-vllm-cpp
Qwen3.6-35B-A3B NVFP4 on vllm.cpp with MTP speculative decoding enabled. Image input is implemented in the engine but is not token-gated against vLLM yet, so the vision usecase on this entry is experimental. The draft head ships inside the checkpoint's own mtp.* tensors, so there is no second model to download. On this model the speculative path is token-exact against speculation-off on both the synchronous and asynchronous schedulers. Same weights as qwen3.6-35b-a3b-nvfp4-vllm-cpp; install that entry instead if you would rather not spend the extra memory on speculative state. MTP is depth 1 by construction on this engine.

Repository: localaiLicense: apache-2.0

qwen3-coder-30b-a3b-vllm-cpp
Qwen3-Coder-30B-A3B on vllm.cpp: a coding and agentic-tool-use model, 30B total parameters with about 3B active per token, gated token-exact against vLLM on this engine. The tool-call parser is named explicitly rather than auto-detected, and that matters here. Qwen3-Coder's tool dialect is byte-identical on the wire to another family's, so template sniffing cannot separate the two and would fall back to the wrong parser. With qwen3_coder named, tool calls arrive as real tool_calls on the OpenAI response. This is the bf16 checkpoint, roughly 57 GB of weights, which is what the engine was gated on. Being bf16 rather than NVFP4 it does not need Blackwell on its own account, but LocalAI's CUDA images for this backend are currently built for Blackwell-family GPUs only, so on an older card use the CPU build.

Repository: localaiLicense: apache-2.0

qwen3-4b-vllm-cpp
Qwen3-4B on vllm.cpp, in bf16. The small end of the engine's gated dense family, which reaches parity with vLLM on every axis at concurrency 1. bf16 rather than NVFP4 on purpose: this is the entry that runs where the flagship NVFP4 checkpoints cannot, including Apple Silicon via Metal, Vulkan and plain CPU. Roughly 8 GB of weights, plus about 4.5 GB of KV cache at the context configured here. Tool calling and the thinking split are parsed inside the engine.

Repository: localaiLicense: apache-2.0

qwen3-0.6b-vllm-cpp
Qwen3-0.6B on vllm.cpp, in bf16. Roughly 1.4 GB of weights, which makes it the cheapest way to confirm a vllm-cpp install actually serves before committing disk and memory to one of the large checkpoints. It runs anywhere the backend does, CPU included, and it is a real chat model rather than a stub, so tool calling and the thinking split can be exercised on it too.

Repository: localaiLicense: apache-2.0

minimax-h3-fl2va-q4
MiniMax-H3 served by vllm.cpp, LocalAI's own C++ port of vLLM. It generates video AND audio jointly from a text prompt, so a clip comes back as an MP4 with a real soundtrack rather than a silent render: ask for speech in the prompt and the model lip-syncs it. This is the Q4_K_M quantisation of the FL2VA partition, which serves text-to-video (t2va) and first/last-frame conditioning (fl2va). Reference conditioning (ref2va) is a different checkpoint and is refused by this one. Roughly 40 GB of weights across five files, plus the two VAE configs that carry the latent statistics. The default canvas is 1344x768 at 124 frames and 24 fps, about 5.2 seconds. Generation is slow — measured at roughly 176 s per denoise step at that canvas on a 20-SM device, so the 50-step default is a multi-hour job. Muxing the finished frames needs ffmpeg on the host.

Repository: localaiLicense: other

minimax-h3-ref2va-q4
MiniMax-H3 served by vllm.cpp, LocalAI's own C++ port of vLLM. It generates video AND audio jointly, so a clip comes back as an MP4 with a real soundtrack rather than a silent render. This is the Q4_K_M quantisation of the Ref2VA partition, the one that takes REFERENCE conditioning: a reference image, a reference clip, or reference audio, prepended as their own blocks so the subject or style carries into the generated video. For plain text-to-video or first/last-frame conditioning use minimax-h3-fl2va-q4 instead - the two partitions are separate checkpoints and each refuses the other's tasks. Use this Q4_K_M build, NOT the NVFP4 Ref2VA weights: NVFP4 renders a multicolour patch grid, and it took three investigations upstream to establish that the fault is the quantisation rather than the reference path. On Q4_K_M the same code renders coherently. Roughly 40 GB of weights across five files, plus the two VAE configs that carry the latent statistics. The default canvas is 1344x768 at 124 frames and 24 fps, about 5.2 seconds. Generation is slow - roughly 176 s per denoise step at that canvas on a 20-SM device, so the 50-step default is a multi-hour job. Muxing the finished frames needs ffmpeg on the host.

Repository: localaiLicense: other

laya-vllm-cpp
Laya is a multilingual, non-autoregressive System 1 decision model. Given a state (text, email, ticket, or JSON) and typed questions, it returns typed answers with mathematically calibrated probabilities in a single forward pass. It never generates text, so there is nothing to parse and nothing to hallucinate. In LocalAI, serve via POST /v1/systemone with this model. The vllm.cpp engine runs the full decision pipeline (choice, noul, score question types) through the vllm_decide C ABI, returning the complete kev-compatible JSON response. ModernBERT-large backbone, 421M params, 512-token context. F16 weights, ~804 MB.

Repository: localaiLicense: apache-2.0

gliner25-decide-vllm-cpp
GLiNER2.5-Decide is a DeBERTa-v3-large encoder with a classification head that answers typed decision questions over a state text in one forward pass. It never generates text, so there is nothing to parse. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp engine runs the decision pipeline (choice, noul and score question types) through the vllm_decide C ABI. This is the decision model, not the zero-shot NER model: use the gliner2.5 entry for entity extraction. F32 weights, about 2 GB. The weights are pinned to a revision so the entry keeps serving the checkpoint it was checked against.

Repository: localaiLicense: apache-2.0

tev1-4b-vllm-cpp
Tev1-4B-experimental is an experimental decision model from Together AI: a supervised fine-tune of Qwen3.5-4B that picks one option letter for a state, a question and 2 to 24 labeled options. It keeps the standard next-token head, so it is autoregressive, unlike Laya or GLiNER2.5-Decide. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp engine scores the answer letters of each choice, noul and score question through the vllm_decide C ABI and returns probabilities with an entropy confidence, as Ollama does for tev1. A choice or score question accepts at most 24 options (Ollama allows 26) and every option needs a nonempty description. The published config.json names Qwen3_5ForConditionalGeneration, so this entry sets hf_overrides to load it as Tev1Model without editing the download. Checked against transformers BF16 on CPU over seven questions: 7/7 answers equal, largest probability difference 0.0004. The decision route is verified on CPU only; GPU serving has not been measured. BF16 weights, about 9.3GB, pinned to a revision. The fine-tune license is still being finalized by Together AI (base model Apache-2.0).

Repository: localai

tev1-0.8b-vllm-cpp
Tev1-0.8B-experimental is an experimental decision model from Together AI: a supervised fine-tune of Qwen3.5-0.8B that picks one option letter for a state, a question and 2 to 24 labeled options. It keeps the standard next-token head, so it is autoregressive, unlike Laya or GLiNER2.5-Decide. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp engine scores the answer letters of each choice, noul and score question through the vllm_decide C ABI and returns probabilities with an entropy confidence, as Ollama does for tev1. A choice or score question accepts at most 24 options (Ollama allows 26) and every option needs a nonempty description. The published config.json names Qwen3_5ForConditionalGeneration, so this entry sets hf_overrides to load it as Tev1Model without editing the download. Checked against transformers BF16 on CPU over seven questions: 6/7 answers equal, the miss a near tie (0.453 against 0.514 in transformers, 0.4845 each here), largest probability difference 0.031. The decision route is verified on CPU only; GPU serving has not been measured. BF16 weights, about 1.8GB, pinned to a revision. The fine-tune license is still being finalized by Together AI (base model Apache-2.0).

Repository: localai

kev-0.8b-vllm-cpp
kev is a System 1 decision model by Jared Palmer. It answers typed choice, noul and score questions about a text state with one scoring pass per question. It does not generate text. The model is a frozen Qwen3.5-0.8B-Base backbone, a rank-16 LoRA adapter and a PointerHead readout. This entry installs a converted redistribution of jaredpalmer/kev-0.8b: the LoRA is merged into the BF16 backbone, the head is stored as head.safetensors, and config.json names the KevModel architecture. The checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers, vLLM and llama.cpp cannot load it. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp project records PointerHead golden-vector tests (25 cases) and a 5-case end-to-end comparison against the kev reference server as equal. The upload itself was smoke-tested with one request on CPU; there is no accuracy benchmark and no GPU run. The entry sets a 2048-token context and an explicit KV pool, because the default 4096-token context does not fit the default CPU KV pool and the load fails. BF16 weights, about 1.53 GB, pinned to a revision.

Repository: localaiLicense: apache-2.0

nimble-9b-vllm-cpp
Nimble is a decision model from Bespoke Labs (the model Ollama serves as nimble). For each question it runs one forward pass and reads the logits of the answer letters at the last prompt position. It answers typed choice, noul and score questions about a text state and does not generate text. This entry installs a converted redistribution of bespokelabs/Bespoke-Nimble-9B: the LoRA adapter is merged into its Qwen3.5-9B base in BF16 and config.json names the NimbleModel architecture. The checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers and vLLM cannot load it. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp project compared this converted directory on CPU against the Bespoke authors' own code over seven questions: 7 of 7 answers equal, largest probability difference 0.0054. This entry was installed and served through LocalAI on CPU and gave the answers shown in the model card example. There is no accuracy benchmark and no GPU (CUDA, ROCm, Metal) run. The engine refuses fields with more than 26 choices (upstream allows 255). The entry sets an 8192-token context (Nimble's own prompt limit) and a KV pool that holds four such sequences. BF16 weights, about 19.3 GB, pinned to a revision. On CPU the model needs about 20 GB of free RAM (measured peak 18.4 GB resident). On a 20-thread CPU a request with three questions (988 prompt tokens) took 30 seconds warm.

Repository: localaiLicense: apache-2.0

clm-v0.1-8b-vllm-cpp
CLM is a bi-encoder decision model from Contrastive-LM. A frozen Qwen3-8B backbone encodes the state and each candidate answer separately, two MLP heads project them into a 512-dimensional space, and the answer distribution is a softmax over the cosine similarities at scale 100. It answers typed choice, noul and score questions and does not generate text. This entry installs a converted redistribution: the CLM-v0.1-8B heads as head.safetensors next to the unchanged Qwen3-8B backbone and tokenizer, with config.json naming the ClmModel architecture. The checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers, vLLM and llama.cpp cannot load it. In LocalAI, serve it via POST /v1/systemone. Put the question in instructions: the state head reads the state followed by the instructions. The vllm.cpp project compared this checkpoint on CPU with the reference code over eight questions: 8 of 8 answers agree, largest probability difference 0.029 against the reference in bf16 (transformers stood in for the reference's GPU encoder). This entry was installed and served through LocalAI on CPU and gave the model card's example answer (person, 0.950). There is no accuracy benchmark and no GPU run. The entry sets a 4096-token context and a KV pool that holds four such sequences. BF16 backbone with F32 heads, about 16.5 GB, pinned to a revision. On CPU the model needs about 19 GB of free RAM (measured peak 17.9 GB resident).

Repository: localaiLicense: apache-2.0

Page 1