tiny Jev-like family of decision models built on top of Qwen3 you can train and run on your own
Kev trains LoRA adapters plus a pointer readout head on Qwen3 (0.6B/4B/8B) to answer typed yes/no, choice, and score questions over a document in one causal prefill pass, returning calibrated probabilities rather than text. Ships a FastAPI /v1/systemone server, frozen eval suites, and a Next.js playground.
Use Cases
Calibrated yes/no, choice, and score decisions over a documentSupport ticket routing and department classificationCustomer frustration and urgency scoringLLM fine-tuning with LoRA adapters for classification headsLocal/on-device decision model serving on a laptopBenchmarking decision models against frozen eval suitesChess position evaluation and move selectionInterpreting LLM probabilities instead of generating proseTypeSafe/System One API-compatible local model endpoint
Built With
- Language
- Python
- Frameworks
- PyTorch · Hugging Face Transformers · PEFT · FastAPI · Uvicorn · Pydantic · Next.js · React · Modal · scikit-learn · accelerate · uv
Tags
decision-model · calibration · lora · qwen3 · inference-engine · serving · evals · benchmarks · probabilities · block-causal-mask · typesafe · fastapi · pytorch · playground · chess · text-classification