Vibe Coding Discover

Use Cases

93 use cases for “Inference”

720p 30fps Fast Video Inference On Single Or Multi-gpu

1 project

Accelerate Llm Inference Throughput And Latency

1 project

Accelerating Deep Learning Inference Across Cpu, Gpu, And Npu

1 project

Accelerating Tts Inference With Vllm Or Tensorrt-llm

1 project

Add A Standalone KV Cache Daemon To Existing Inference Engines

1 project

Add Local Llm Inference To Existing Sdk-based Applications

1 project

Analyze Llm Inference Stack Performance With Aiperf

1 project

Batch Inference Over Software Tasks

1 project

Benchmark Disk-streaming Vs Resident Inference Performance

1 project

Benchmark Inference Latency And Throughput

1 project

Benchmark Inference Performance Across Optimized Models

1 project

Benchmark Inference Speed And Concurrency

1 project

Benchmarking Inference Latency And Throughput Vs Upstream Pytorch Ports

1 project

Benchmarking Inference Throughput And Latency

1 project

Build Lightweight C/c++ Inference Into Apps And Edge Devices

1 project

Build Python Model Pipelines For Inference

1 project

Cli Inference For Batch Audio Conversion

1 project

Cpu Inference Of Sharded Models

1 project

Cpu-only Inference With No Gpu

1 project

Cpu-only Local Llm Inference Without Gpu Or Blas

1 project

Custom-metrics Autoscaling For Inference Services

1 project

Cut Inference Cost For Coding Agents Without Changing Agent Code

1 project

Deploying A Tts Inference Server Via Grpc/fastapi/docker

1 project

Disk-backed Local Cluster Inference Across Machines

1 project

Distributed Inference And Fine-tuning Across Macs

1 project

Distributed Inference Over Rpc Backend

1 project

Distributed Multi-gpu/multi-node Inference

1 project

Distributed Multi-mac Inference Over Ring/thunderbolt

1 project

Distributed Multi-node Inference Cluster

1 project

Edge Inference

1 project

Export Mask Decoder To Onnx For In-browser Inference

1 project

Fast Low-latency Per-decision Inference (e.g. Game Ai)

1 project

Gpu-accelerated Local Inference With Nvidia Cuda

1 project

High-throughput Batch Image Generation With Vllm Offline Inference

1 project

High-throughput Llm Inference And Serving

1 project

Hybrid Cpu+gpu Inference For Models Larger Than Vram

1 project

Improve Moe Inference Throughput

1 project

Indicate Model Inference Or Loading

1 project

Inference On Apple Silicon Macs

1 project

Inference Routing To Self-hosted Models On Kubernetes

1 project

Integrate Diffusion Inference Into C/c++ Applications

1 project

Local Inference Without Sending Data To A Hosted Api

1 project

Local Offline Llm Inference With Ollama/vllm/lm Studio

1 project

Local Private Llm Inference Without Api Providers

1 project

Local-first Inference With Llama.cpp

1 project

Long-context Inference With Compressed Mla KV State

1 project

Low-latency Chat/completion Inference At Scale

1 project

Machine-learning Inference

1 project

Memory-efficient Inference On Long Documents Via Compact Kv-cache Merging

1 project

Multi-node Distributed Inference With Tensor/expert Parallelism

1 project

Multimodal Image/video/audio Chat Inference

1 project

Neural Vocoder Inference

1 project

Offline Inference With No Api Keys Or Accounts

1 project

Offline Inference Without Cloud Api Or Output Tokens

1 project

Offline On-device Inference Via Llama.cpp With Lora Adapters

1 project

Offline Or Air-gapped Inference

1 project

On-device Inference Via Llama.cpp-omni

1 project

Point Claude, Codex, And Opencode At Custom Inference Endpoints Or Byo Models

1 project

Prefill-decode Disaggregated And Multi-node Inference

1 project

Prefill/decode Disaggregation And Dp-aware Routing For Inference Engines

1 project

Private On-prem Multi-modal Inference

1 project

Quantize Models For Faster Inference And Smaller Checkpoints

1 project

Quantized Inference (fp4/fp8/int4/awq/gptq)

1 project

RL Rollout Infrastructure Spanning Agent, Inference And Training

1 project

Recursive Inference Over Inputs Larger Than The Model Context Window

1 project

Route Tasks To Best-fit Models Through An Optional Inference Router

1 project

Run 70b Llm Inference On A Single 4gb Gpu

1 project

Run A First Local Llm Inference With Ollama

1 project

Run AI Inference Close To End Users

1 project

Run Bedrock Batch Inference Jobs Via Step Functions

1 project

Run Inference Through Cli Tools Like Claude Or Gemini

1 project

Run Inference/training Through A Local Gradio Webui

1 project

Run Local Image Inference On Mac Via Sd.cpp

1 project

Run Local Llm Inference For Music-production Workflows

1 project

Run Local On-device Llm Inference

1 project

Run Local Video Inference Via Wan2gp Server

1 project

Run On-device Inference With Local Models

1 project

Run Private/offline Inference Without Sending Data To Cloud Apis

1 project

Run Quantized Llm Inference In The Browser Via Onnx + Wasm

1 project

Run Uncensored 27b Inference Locally On Apple Silicon

1 project

Running Fully Local/private Inference Via Privacy Mode Or Byok Providers

1 project

Running Llm Inference Locally With Optimized Runtime Backends

1 project

Self-host An Openai-compatible Inference Server

1 project

Serve An Openai-compatible Local Inference Server

1 project

Serve Batched State/question/candidate Requests Over An Http Inference Api

1 project

Speed Up Inference 3x With Block-wise Quantization

1 project

Start/stop The Local LM Studio Inference Api Server

1 project

Stream Model Inference Logs To The Terminal

1 project

Study A From-scratch Transformer/moe Inference Engine In Portable C99

1 project

Summarizing Very Large Logs With Budgeted Fan-out Inference

1 project

Sync And Async Inference Sessions Over Websocket

1 project

Unwrapping Lightning Checkpoints For Inference

1 project

Use The Gateway As A Unified Authenticated Proxy For Model Inference Endpoints

1 project