NanoJev
View on GitHubA nano replica of Jev: parallel decisions, dynamic candidates, and an end-to-end training pipeline.
NanoJev is a 0.6B Qwen3-based replica of Jev: a parallel decision model that returns probability distributions over supplied candidate actions with zero output-token decoding. Ships SFT training configs, an 18.7K-question dataset, an HTTP decision service and ViZDoom/Maze/Snake benchmarks.
Use Cases
Train a 0.6B model to output action probability distributions without decoding tokensBenchmark small decision models against LLM baselines on ViZDoom, Maze and SnakeServe batched state/question/candidate requests over an HTTP inference APIBuild multi-task game-playing policies from mixed SFT data splitsReplay recorded agent episodes in browser viewersBatch independent states and candidate paths in one backbone forward pass
Built With
- Language
- Python
- Frameworks
- PyTorch · Hugging Face Transformers · Hugging Face Hub · Qwen3 · ViZDoom · Vercel AI SDK · Triton
Tags
decision-model · parallel-decoding · game-agents · ViZDoom · reinforcement-learning · SFT · probability-heads · training-pipeline · inference-server · Qwen3 · small-language-model · benchmarks · maze · snake · action-selection · replay-viewer