Qwen 3.5 0.8B parameter model quantized for llama-cpp backend. Supports chat interactions and multimodal image-text inputs.
Links
Tags
Repository: localai
Tev1-0.8B-experimental is an experimental decision model from Together AI: a supervised fine-tune of Qwen3.5-0.8B that picks one option letter for a state, a question and 2 to 24 labeled options. It keeps the standard next-token head, so it is autoregressive, unlike Laya or GLiNER2.5-Decide. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp engine scores the answer letters of each choice, noul and score question through the vllm_decide C ABI and returns probabilities with an entropy confidence, as Ollama does for tev1. A choice or score question accepts at most 24 options (Ollama allows 26) and every option needs a nonempty description. The published config.json names Qwen3_5ForConditionalGeneration, so this entry sets hf_overrides to load it as Tev1Model without editing the download. Checked against transformers BF16 on CPU over seven questions: 6/7 answers equal, the miss a near tie (0.453 against 0.514 in transformers, 0.4845 each here), largest probability difference 0.031. The decision route is verified on CPU only; GPU serving has not been measured. BF16 weights, about 1.8GB, pinned to a revision. The fine-tune license is still being finalized by Together AI (base model Apache-2.0).
Links
Tags
Repository: localaiLicense: apache-2.0
kev is a System 1 decision model by Jared Palmer. It answers typed choice, noul and score questions about a text state with one scoring pass per question. It does not generate text. The model is a frozen Qwen3.5-0.8B-Base backbone, a rank-16 LoRA adapter and a PointerHead readout. This entry installs a converted redistribution of jaredpalmer/kev-0.8b: the LoRA is merged into the BF16 backbone, the head is stored as head.safetensors, and config.json names the KevModel architecture. The checkpoint only works with vllm.cpp (the vllm-cpp backend); transformers, vLLM and llama.cpp cannot load it. In LocalAI, serve it via POST /v1/systemone. The vllm.cpp project records PointerHead golden-vector tests (25 cases) and a 5-case end-to-end comparison against the kev reference server as equal. The upload itself was smoke-tested with one request on CPU; there is no accuracy benchmark and no GPU run. The entry sets a 2048-token context and an explicit KV pool, because the default 4096-token context does not fit the default CPU KV pool and the load fails. BF16 weights, about 1.53 GB, pinned to a revision.
Links
Tags