Model Gallery

8 models from 1 repositories

Filter by type:

Filter by tags:

grug-12b
Grug 12B is kai-os's compact-reasoning fine-tune of Gemma 4 12B IT. It targets shorter, denser reasoning traces while preserving constraints, branching decisions, edge cases, and final-answer checks. This entry uses Bartowski's Q4_K_M quantization and includes the multimodal projector for Gemma 4 image inputs. The model is experimental and its reported evaluation is a small local math proxy rather than a broad benchmark. Review the upstream model card's dataset provenance and `other` license before commercial or sensitive use.

Repository: localaiLicense: other

laya-vllm-cpp
Laya is a multilingual, non-autoregressive System 1 decision model. Given a state (text, email, ticket, or JSON) and typed questions, it returns typed answers with mathematically calibrated probabilities in a single forward pass. It never generates text, so there is nothing to parse and nothing to hallucinate. In LocalAI, serve via POST /v1/systemone with this model. The vllm.cpp engine runs the full decision pipeline (choice, noul, score question types) through the vllm_decide C ABI, returning the complete kev-compatible JSON response. ModernBERT-large backbone, 421M params, 512-token context. F16 weights, ~804 MB.

Repository: localaiLicense: apache-2.0

gliner25-decide-vllm-cpp
GLiNER2.5-Decide is a DeBERTa-v3-large encoder with a classification head that answers typed decision questions over a state text in one forward pass. It never generates text, so there is nothing to parse. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp engine runs the decision pipeline (choice, noul and score question types) through the vllm_decide C ABI. This is the decision model, not the zero-shot NER model: use the gliner2.5 entry for entity extraction. F32 weights, about 2 GB. The weights are pinned to a revision so the entry keeps serving the checkpoint it was checked against.

Repository: localaiLicense: apache-2.0

tev1-4b-vllm-cpp
Tev1-4B-experimental is an experimental decision model from Together AI: a supervised fine-tune of Qwen3.5-4B that picks one option letter for a state, a question and 2 to 24 labeled options. It keeps the standard next-token head, so it is autoregressive, unlike Laya or GLiNER2.5-Decide. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp engine scores the answer letters of each choice, noul and score question through the vllm_decide C ABI and returns probabilities with an entropy confidence, as Ollama does for tev1. A choice or score question accepts at most 24 options (Ollama allows 26) and every option needs a nonempty description. The published config.json names Qwen3_5ForConditionalGeneration, so this entry sets hf_overrides to load it as Tev1Model without editing the download. Checked against transformers BF16 on CPU over seven questions: 7/7 answers equal, largest probability difference 0.0004. The decision route is verified on CPU only; GPU serving has not been measured. BF16 weights, about 9.3GB, pinned to a revision. The fine-tune license is still being finalized by Together AI (base model Apache-2.0).

Repository: localai

tev1-0.8b-vllm-cpp
Tev1-0.8B-experimental is an experimental decision model from Together AI: a supervised fine-tune of Qwen3.5-0.8B that picks one option letter for a state, a question and 2 to 24 labeled options. It keeps the standard next-token head, so it is autoregressive, unlike Laya or GLiNER2.5-Decide. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp engine scores the answer letters of each choice, noul and score question through the vllm_decide C ABI and returns probabilities with an entropy confidence, as Ollama does for tev1. A choice or score question accepts at most 24 options (Ollama allows 26) and every option needs a nonempty description. The published config.json names Qwen3_5ForConditionalGeneration, so this entry sets hf_overrides to load it as Tev1Model without editing the download. Checked against transformers BF16 on CPU over seven questions: 6/7 answers equal, the miss a near tie (0.453 against 0.514 in transformers, 0.4845 each here), largest probability difference 0.031. The decision route is verified on CPU only; GPU serving has not been measured. BF16 weights, about 1.8GB, pinned to a revision. The fine-tune license is still being finalized by Together AI (base model Apache-2.0).

Repository: localai

kev-0.8b-vllm-cpp
kev is a System 1 decision model by Jared Palmer. It answers typed choice, noul and score questions about a text state with one scoring pass per question. It does not generate text. The model is a frozen Qwen3.5-0.8B-Base backbone, a rank-16 LoRA adapter and a PointerHead readout. This entry installs a converted redistribution of jaredpalmer/kev-0.8b: the LoRA is merged into the BF16 backbone, the head is stored as head.safetensors, and config.json names the KevModel architecture. The checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers, vLLM and llama.cpp cannot load it. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp project records PointerHead golden-vector tests (25 cases) and a 5-case end-to-end comparison against the kev reference server as equal. The upload itself was smoke-tested with one request on CPU; there is no accuracy benchmark and no GPU run. The entry sets a 2048-token context and an explicit KV pool, because the default 4096-token context does not fit the default CPU KV pool and the load fails. BF16 weights, about 1.53 GB, pinned to a revision.

Repository: localaiLicense: apache-2.0

nimble-9b-vllm-cpp
Nimble is a decision model from Bespoke Labs (the model Ollama serves as nimble). For each question it runs one forward pass and reads the logits of the answer letters at the last prompt position. It answers typed choice, noul and score questions about a text state and does not generate text. This entry installs a converted redistribution of bespokelabs/Bespoke-Nimble-9B: the LoRA adapter is merged into its Qwen3.5-9B base in BF16 and config.json names the NimbleModel architecture. The checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers and vLLM cannot load it. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp project compared this converted directory on CPU against the Bespoke authors' own code over seven questions: 7 of 7 answers equal, largest probability difference 0.0054. This entry was installed and served through LocalAI on CPU and gave the answers shown in the model card example. There is no accuracy benchmark and no GPU (CUDA, ROCm, Metal) run. The engine refuses fields with more than 26 choices (upstream allows 255). The entry sets an 8192-token context (Nimble's own prompt limit) and a KV pool that holds four such sequences. BF16 weights, about 19.3 GB, pinned to a revision. On CPU the model needs about 20 GB of free RAM (measured peak 18.4 GB resident). On a 20-thread CPU a request with three questions (988 prompt tokens) took 30 seconds warm.

Repository: localaiLicense: apache-2.0

clm-v0.1-8b-vllm-cpp
CLM is a bi-encoder decision model from Contrastive-LM. A frozen Qwen3-8B backbone encodes the state and each candidate answer separately, two MLP heads project them into a 512-dimensional space, and the answer distribution is a softmax over the cosine similarities at scale 100. It answers typed choice, noul and score questions and does not generate text. This entry installs a converted redistribution: the CLM-v0.1-8B heads as head.safetensors next to the unchanged Qwen3-8B backbone and tokenizer, with config.json naming the ClmModel architecture. The checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers, vLLM and llama.cpp cannot load it. In LocalAI, serve it via POST /v1/systemone. Put the question in instructions: the state head reads the state followed by the instructions. The vllm.cpp project compared this checkpoint on CPU with the reference code over eight questions: 8 of 8 answers agree, largest probability difference 0.029 against the reference in bf16 (transformers stood in for the reference's GPU encoder). This entry was installed and served through LocalAI on CPU and gave the model card's example answer (person, 0.950). There is no accuracy benchmark and no GPU run. The entry sets a 4096-token context and a KV pool that holds four such sequences. BF16 backbone with F32 heads, about 16.5 GB, pinned to a revision. On CPU the model needs about 19 GB of free RAM (measured peak 17.9 GB resident).

Repository: localaiLicense: apache-2.0