collabosm
View on GitHubRun Qwen3.8-Flash-Next (125B-A6B MoE) on ONE Colab A100-80GB High-RAM: pinned runtime, weights from HF, OpenAI-compatible endpoint, measured 2.8k-3.9k t/s prefill / 90-97 t/s decode, plus a verified pinned-RAM KV tier.
Scripts to deploy Qwen3.8-Flash-Next on a high-RAM Colab A100 and expose it through an authenticated OpenAI-compatible API. Includes runtime and model setup, session recovery, performance measurements, and KV-cache management.
Use Cases
Serve Qwen3.8-Flash-Next from a Colab A100Expose a model through an OpenAI-compatible APIBenchmark inference speed and concurrencyResume cached conversations using pinned host RAM
Built With
- Language
- Python
- Frameworks
- ExLlamaV3 · Hugging Face
Tags
LLM inference · model serving · OpenAI-compatible API · Google Colab · A100 · ExLlamaV3 · KV cache · Qwen