Category
AI Frameworks
Core frameworks for building LLM applications.
| Project name | Stars | Category | Language | Tags | Summary | View Project |
|---|---|---|---|---|---|---|
| tensorflow | ★ 200K | AI Frameworks | C++ | machine-learning · deep-learning · neural-networks · distributed-training · GPU · model-training · model-deployment | An end-to-end machine-learning framework with Python and C++ APIs for building, training, and deploying models across CPUs, GPUs, and other devices. | View Project → |
| ollama | ★ 181K | AI Frameworks | Go | llm · local-llm · inference · model-runner · llama.cpp · gguf · quantization · cli · rest-api · self-hosted · go · model-management · openai-compatible · embeddings · offline-ai · developer-tools | Ollama is a Go-based local LLM runtime that downloads, runs, and serves open models (Gemma, Qwen, DeepSeek, Llama) via a CLI and REST API on port 11434. It powers local inference for coding agents, chat UIs, and RAG apps, and supports custom Modelfiles and imports. | View Project → |
| transformers | ★ 166K | AI Frameworks | Python | transformers · model-definition · pretrained-models · llm · multimodal · nlp · computer-vision · audio · speech-recognition · inference · training · fine-tuning · pytorch · model-hub · vlm · pipeline-api | Hugging Face Transformers is the model-definition framework for state-of-the-art text, vision, audio, video and multimodal models. It offers a unified Pipeline/Trainer API over 1M+ Hub checkpoints and feeds model definitions to vLLM, llama.cpp and training stacks. | View Project → |
| dify | ★ 157K | AI Frameworks | TypeScript | llm · agents · agentic-workflow · rag · low-code · no-code · workflow-builder · mcp · llmops · observability · prompt-ide · self-hosted · multi-model · baas · tool-calling · orchestration | Dify is an open-source LLM app platform combining a visual agentic workflow builder, RAG pipeline, prompt IDE, agent runtime, and LLMOps. It supports hundreds of models, MCP servers, and tools, and self-hosts via Docker with APIs for integration. | View Project → |
| langflow | ★ 155K | AI Frameworks | Python | visual-builder · low-code · multi-agent · workflows · orchestration · llm · mcp-server · rag · python · react-flow · api-deployment · playground · vector-databases · observability · self-hosted · no-code | Langflow is an open-source visual platform for building and deploying LLM agents and workflows. It combines a drag-and-drop React Flow canvas with Python component customization, then ships each flow as a REST API, an MCP server, or exportable JSON. | View Project → |
| langchain | ★ 147K | AI Frameworks | Python | agents · llm · rag · multi-agent · tool-calling · orchestration · prompt-templates · chat-models · embeddings · vector-stores · structured-output · streaming · integrations · model-interoperability · python · typescript | LangChain is the Python/JS framework for building LLM apps and agents: a standard interface for chat models, embeddings, vector stores, tools, and retrievers, plus chains, structured output, and streaming. Paired with LangGraph for controllable agent workflows. | View Project → |
| llama.cpp | ★ 129K | AI Frameworks | C++ | llm-inference · gguf · quantization · local-llm · cpu-inference · gpu-inference · cpp · openai-compatible-api · cuda · metal · vulkan · multimodal · edge-inference · no-dependencies · inference-engine · model-conversion | llama.cpp is a dependency-free C/C++ LLM and VLM inference engine built on ggml. It runs quantized GGUF models on CPU, CUDA, Metal, Vulkan and many other backends, and ships CLI, server, web UI and OpenAI-compatible API tools. | View Project → |
| llama.cpp | ★ 129K | AI Frameworks | C++ | C++ · llama.cpp · Inference | High-performance C/C++ inference for local LLMs across CPUs, GPUs, and devices. | View Project → |
| vllm | ★ 92K | AI Frameworks | Python | inference-engine · llm-serving · paged-attention · continuous-batching · quantization · speculative-decoding · moe · openai-compatible-api · kv-cache · distributed-inference · lora · structured-output · cuda · rocm · tpu · throughput | vLLM is a high-throughput, memory-efficient LLM inference and serving engine built on PagedAttention, continuous batching, and optimized CUDA/ROCm kernels. It exposes an OpenAI-compatible server with quantization, LoRA, speculative decoding, and disaggregated serving across 200+ model architectures. | View Project → |
| unsloth | ★ 77K | AI Frameworks | Python | fine-tuning · llm · lora · qlora · grpo · reinforcement-learning · gguf · quantization · inference · local-llm · diffusion · tts · training · desktop-app · openai-api · mcp | Unsloth is a Python training and inference stack plus desktop/web UI for running, fine-tuning and RL-training LLMs and diffusion models 2x faster with less VRAM. Supports LoRA/QLoRA/full fine-tuning, GGUF/MLX export, local OpenAI-compatible serving, RAG and agent hooks. Apache-2.0. | View Project → |
| annotated_deep_learning_paper_implementations | ★ 67K | AI Frameworks | Python | transformers · pytorch · deep-learning · attention · llm · diffusion · gan · reinforcement-learning · optimizers · lora · education · tutorial · neural-networks · literate-programming · vit · stable-diffusion | LabML's annotated deep learning library: 60+ readable PyTorch implementations of papers with side-by-side notes, covering transformers (GPT, ViT, XL), LoRA, diffusion/Stable Diffusion, GANs, optimizers, and RL. Best as a learning and reference resource rather than a production framework. | View Project → |
| nanoGPT | ★ 63K | AI Frameworks | Python | gpt · llm-training · finetuning · transformer · pytorch · gpt-2 · text-generation · minimal · educational · ddp · character-level · openwebtext · shakespeare · reproduction · deprecated · nanochat | Minimal, readable ~300-line PyTorch repo for training and finetuning GPT-2-scale models; reproduces GPT-2 124M on OpenWebText. Author marks it deprecated in favor of the newer nanochat, but it remains a compact hacking/education base. | View Project → |
| litellm | ★ 59K | AI Frameworks | Python | ai-gateway · llm-gateway · llmops · openai-compatible · proxy-server · cost-tracking · guardrails · load-balancing · mcp-gateway · a2a · multi-provider · python-sdk · rust · observability · virtual-keys · rate-limiting | LiteLLM is an open-source AI gateway and Python SDK that calls 100+ LLM providers (OpenAI, Anthropic, Bedrock, Vertex, vLLM, Ollama) in OpenAI format. The proxy adds virtual keys, spend tracking, guardrails, load balancing, logging, plus MCP tool and A2A agent routing. | View Project → |
| llama_index | ★ 52K | AI Frameworks | Python | rag · llm · agents · multi-agent · workflows · vector-database · retrieval · data-connectors · document-parsing · indexing · embeddings · python · llm-orchestration · query-engine · structured-extraction | LlamaIndex is a Python framework for building LLM apps: data connectors, indices, retrievers and query engines for RAG, plus agent/workflow abstractions. 300+ integrations let you swap LLMs, embeddings and vector stores; the team now also pushes the hosted LlamaParse document platform. | View Project → |
| LocalAI | ★ 49K | AI Frameworks | Go | llm-inference · local-inference · openai-compatible-api · self-hosted · multimodal · cpu-inference · gpu-acceleration · agents · mcp · rag · text-to-speech · speech-recognition · image-generation · video-generation · distributed-inference · docker | LocalAI is a self-hosted, OpenAI/Anthropic-compatible inference engine that runs LLMs, vision, voice, image and video models on CPU or GPU. Backends like llama.cpp, vLLM and whisper.cpp are pulled on demand, and it also ships agents, RAG and MCP support. | View Project → |
| agno | ★ 42K | AI Frameworks | Python | agents · multi-agent · agent-runtime · agentos · python · llm · mcp · rag · memory · observability · scheduling · human-in-the-loop · self-hosted · rest-api · toolkits · auth | Agno is a Python framework plus AgentOS runtime for building, serving, and managing agent platforms. It adds REST/SSE endpoints, storage, memory, RAG knowledge, 100+ toolkits, human approval, scheduling, RBAC and OpenTelemetry observability, with Docker and cloud deploy templates. | View Project → |
| langgraph | ★ 42K | AI Frameworks | Python | agents · multi-agent · orchestration · stateful-workflows · durable-execution · human-in-the-loop · memory · graph · python · langchain · observability · langsmith · checkpointing · workflow | LangGraph is a low-level Python orchestration framework for building stateful, long-running agents as graphs. It adds durable execution, checkpointing, human-in-the-loop interrupts, and memory, with optional LangSmith tracing and deployment. | View Project → |
| dspy | ★ 38K | AI Frameworks | Python | dspy · prompt-optimization · llm-programming · declarative · pipelines · compiler · teleprompters · few-shot · self-improving · gepa · signatures · modules · evaluation · python | DSPy is a Python framework for programming rather than prompting LLMs: you write declarative modules and signatures, then compilers/optimizers (e.g. GEPA, MIPRO) tune prompts and weights against your metric. Useful for building and optimizing RAG pipelines, classifiers, and agent loops. | View Project → |
| colibri | ★ 38K | AI Frameworks | C | inference-engine · moe · local-llm · c · zero-dependencies · quantization · int4 · expert-streaming · cpu-inference · gpu-backends · kv-cache · speculative-decoding · disk-offloading · llm-serving · nuive-streaming · memory-hierarchy | Colibrì is a pure-C, zero-dependency inference engine that runs frontier MoE models (GLM-5.x 744B, Kimi K3 2.8T, DeepSeek V4, Qwen3.x) on consumer hardware by streaming disk-resident experts into a unified VRAM/RAM/NVMe tier hierarchy. It offers chat/serve/web front ends plus CPU, CUDA, Metal and Vulkan backends. | View Project → |
| CopilotKit | ★ 37K | AI Frameworks | TypeScript | agent-native · generative-ui · react · angular · vue · react-native · slack · microsoft-teams · ag-ui-protocol · human-in-the-loop · shared-state · chat-ui · typescript · copilot · agent-skills · mcp-apps | CopilotKit is a TypeScript SDK for building agent-native apps: chat UI, generative UI, shared state, and human-in-the-loop across React, Angular, Vue, React Native, Slack, and Teams. It also maintains the AG-UI protocol connecting agent frameworks to frontends. | View Project → |
| sglang | ★ 36K | AI Frameworks | Python | llm-serving · inference-engine · radixattention · prefix-caching · speculative-decoding · quantization · moe · multimodal · vlm · diffusion · distributed-inference · openai-compatible · reinforcement-learning · continuous-batching · paged-attention · high-throughput | SGLang is a high-performance serving framework for LLMs and multimodal models, offering RadixAttention prefix caching, prefill-decode disaggregation, speculative decoding, and quantization over an OpenAI-compatible API across NVIDIA, AMD, TPU, and CPU hardware. | View Project → |
| airllm | ★ 35K | AI Frameworks | Jupyter Notebook | llm-inference · memory-optimization · low-vram · quantization · model-compression · layer-streaming · gpu · macos · apple-silicon · bitsandbytes · lora-finetuning · fp8 · prefetching · pytorch · inference-engine | AirLLM is a Python inference engine that runs very large LLMs (70B up to 671B/2.8T) on a single 4GB GPU without quantization, distillation, or pruning by streaming model layers from disk. Optional 4/8-bit block-wise compression gives ~3x speedup, and it supports macOS/Apple Silicon plus low-VRAM training of 125B models | View Project → |
| VoiceStudio | ★ 35K | AI Frameworks | Python | voice-cloning · text-to-speech · speech-to-text · transcription · dubbing · local-first · tts · audiobook · dictation · voice-design · multilingual · mcp · self-hosted · electron · tauri · python | Fully local, open-source ElevenLabs alternative for voice cloning, voice design, dubbing, dictation, transcription and audiobooks in 646 languages. Ships a desktop app (Electron/Tauri), model manager, local API and an MCP server for agent integration. | View Project → |
| agentscope | ★ 33K | AI Frameworks | Python | multi-agent · agent-framework · react-agent · mcp · rag · realtime · voice-agent · a2a · hitl · sandbox · memory · middleware · multimodal · agent-service · tool-use · python | AgentScope 2.0 is Alibaba Tongyi's Python framework for building and running production agents: ReAct agents, toolkits over Python/MCP/skills, multi-agent teams, realtime voice, sandboxed execution, HITL, memory, plus a FastAPI agent service with web UI and IM channels. | View Project → |
| modular | ★ 30K | AI Frameworks | Mojo | mojo · max · llm-inference · inference-server · openai-compatible · gpu-kernels · model-pipelines · compiler · mlir · python · ai-platform · deployment | Modular's open-source platform hosting MAX (an inference framework with an OpenAI-compatible server and accelerator kernels) and the Mojo language, including its compiler, standard library, and Python model pipelines. | View Project → |
| semantic-kernel | ★ 29K | AI Frameworks | C# | ai-agents · multi-agent · orchestration · llm · sdk · plugins · function-calling · mcp · prompt-templates · vector-db · rag · memory · multimodal · dotnet · python · java | Microsoft's model-agnostic SDK for building AI agents and multi-agent systems in Python, .NET, and Java. Ships plugins/tool calling, memory and vector connectors, planners, and MCP support; now succeeded by Microsoft Agent Framework. | View Project → |
| mastra | ★ 28K | AI Frameworks | TypeScript | typescript · ai-agents · workflows · mcp · rag · evals · observability · memory · llm · model-routing · tool-calling · human-in-the-loop · tts · nextjs · nodejs · chatbots | Mastra is a TypeScript framework for building AI agents and LLM apps: agents, graph workflows, RAG, memory, MCP servers, evals, and observability, with model routing across 40+ providers and React/Next.js/Node integration. | View Project → |
| ai | ★ 27K | AI Frameworks | TypeScript | typescript · ai-sdk · llm · provider-agnostic · agents · tool-calling · streaming · generative-ui · structured-output · embeddings · react · nextjs · vercel · multi-provider · chat-ui · zod | Vercel AI SDK is a provider-agnostic TypeScript toolkit for building AI apps and agents: unified model APIs, streaming text, structured output, tool loops, and framework-agnostic UI hooks for React, Next.js, Svelte and Vue. | View Project → |
| haystack | ★ 27K | AI Frameworks | Python | llm-orchestration · rag · agents · pipelines · retrieval · semantic-search · context-engineering · python · tool-calling · memory · multimodal · evaluation · async · vector-search · multi-agent · mcp | Haystack is deepset's open-source Python framework for building production LLM apps: modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Model- and vendor-agnostic, with built-in components for RAG, semantic search, evaluation, and async streaming. | View Project → |
| laya | ★ 26K | AI Frameworks | Python | decision-model · classification · multilingual · calibrated-probabilities · zero-shot · routing · typed-decisions · single-pass-inference · MCP · LangChain | Python library for fast, calibrated typed decisions over text, including classification, scoring, and yes/no probabilities in a single model pass. It routes requests to English or multilingual checkpoints and supports HTTP, MCP, and LangChain integrations. | View Project → |
| omlx | ★ 22K | AI Frameworks | Python | llm-inference · apple-silicon · mlx · openai-api · continuous-batching · kv-cache · local-llm · macos · menubar-app · vlm · embeddings · reranker · quantization · mcp · self-hosted · ssd-cache | oMLX is an MLX-based LLM inference server for Apple Silicon with continuous batching, tiered hot-RAM/cold-SSD KV cache, multi-model serving, and VLMs/embeddings/rerankers behind an OpenAI-compatible API, managed from a macOS menu bar app and web dashboard. | View Project → |
| onnxruntime | ★ 22K | AI Frameworks | C++ | onnx · inference-engine · model-serving · hardware-acceleration · cross-platform · quantization · graph-optimization · deep-learning · machine-learning · training · edge-deployment · cuda · tensorrt · cpp · python · performance | Cross-platform ML inference and training accelerator. Runs models from PyTorch, TensorFlow, and classical ML libraries via ONNX, applying graph optimizations and hardware acceleration on CPU, GPU, or NPU. Widely used runtime for deploying fast, low-cost model inference. | View Project → |
| peft | ★ 22K | AI Frameworks | Python | peft · lora · fine-tuning · parameter-efficient · adapters · qlora · soft-prompts · ia3 · diffusers · quantization · llm · pytorch · transformers · training · inference · huggingface | Hugging Face PEFT implements parameter-efficient fine-tuning methods (LoRA, QLoRA, IA3, prompt tuning, adapters and more) on top of Transformers, Diffusers, Accelerate and TRL. It lets you train and serve large models with a fraction of the GPU memory and storage, with only small adapter checkpoints. | View Project → |
| pydantic-ai | ★ 20K | AI Frameworks | Python | python · agent-framework · llm · pydantic · type-safety · structured-outputs · tool-calling · multi-agent · mcp · realtime-voice · image-generation · embeddings · evals · durable-execution · dependency-injection · cli | Pydantic AI is a Python AI agent framework from the Pydantic team: a typed agent loop with validated structured outputs, tools, dependency injection, MCP, multi-agent support, realtime voice, image generation, embeddings and evals, with provider-agnostic model switching. | View Project → |
| json-render | ★ 18K | AI Frameworks | TypeScript | generative-ui · json-spec · streaming · ui-generation · llm · zod-schema · shadcn-ui · react · cross-platform · devtools · mcp · remotion · 3d · pdf · email · codegen | Vercel Labs' Generative UI framework: an LLM emits JSON specs constrained to a Zod-typed component catalog, which json-render streams and renders safely across React, Vue, Svelte, Solid, React Native, Next.js, Remotion, PDF, email, 3D and terminal targets. | View Project → |
| rocketride-server | ★ 18K | AI Frameworks | Python | ai-pipeline · llm-workflows · agent-orchestration · rag · mcp · vector-database · cpp-runtime · visual-builder · vscode-extension · python-sdk · typescript-sdk · ocr · ner · pii-anonymization · observability · docker | MIT-licensed AI pipeline engine with a multithreaded C++ runtime and 50+ Python-extensible nodes. Compose LLM, RAG and multi-agent workflows as portable JSON in VS Code or via Python/TypeScript/MCP SDKs, backed by 13+ model providers and 9 vector DBs. | View Project → |
| instructor | ★ 14K | AI Frameworks | Python | Python · Structured Output · Pydantic | Structured LLM outputs validated by Pydantic, with retries and streaming. | View Project → |
| agent-framework | ★ 14K | AI Frameworks | Python | agents · multi-agent · orchestration · workflows · python · dotnet · sdk · middleware · observability · human-in-the-loop · agent-skills · declarative-agents · opentelemetry · hosting · checkpointing · llm-providers | Microsoft's open multi-language SDK for building, orchestrating and hosting production AI agents and multi-agent workflows in Python and .NET. Ships graph workflows, middleware, OpenTelemetry observability, declarative YAML agents and Foundry hosting. | View Project → |
| semantica | ★ 13K | AI Frameworks | Python | knowledge-graph · context-graph · graph-rag · provenance · ai-governance · decision-intelligence · ontology · agent-memory · explainable-ai · semantic-search · reasoning-engine · entity-resolution · mcp · sparql · shacl · python | Semantica is a Python graph-native context and knowledge-graph layer for AI systems: ingest data, build context graphs, and run deterministic graph reasoning with W3C PROV-O provenance. Adds graph RAG, decision audit trails, ontology governance, and MCP/LangChain/CrewAI integrations on top of your existing LLM stack. | View Project → |
| eino | ★ 13K | AI Frameworks | Go | golang · llm-framework · agent-development-kit · multi-agent · orchestration · graph-workflows · streaming · tool-calling · react-agent · human-in-the-loop · interrupt-resume · callbacks · rag · chatmodel · langchain-alternative · byteDance | Eino is ByteDance's Go LLM application framework, modeled on LangChain and Google ADK. It offers component abstractions (ChatModel, Tool, Retriever), a graph/compose orchestration layer with automatic streaming, an agent development kit with ReAct and multi-agent patterns, and interrupt/resume for human-in-the-loop. | View Project → |
| langchain4j | ★ 13K | AI Frameworks | Java | java · jvm · llm · rag · agents · tool-calling · mcp · embeddings · vector-store · chat-memory · prompt-templating · spring-boot · quarkus · embeddings-store · guardrails · ai-orchestration | LangChain4j is an idiomatic Java library for building LLM-powered JVM apps, with a unified API over 20+ model providers and 30+ embedding stores. It covers tool calling (incl. MCP), agents, RAG pipelines, chat memory and prompt templating, with Spring Boot and Quarkus integrations. | View Project → |
| needle | ★ 13K | AI Frameworks | Python | on-device-ai · edge-ai · tinyml · tool-calling · function-calling · llm · embeddings · structured-extraction · quantization · 2-bit · lora · fine-tuning · foundation-model · microcontrollers · wearables · offline-inference | Needle is an 8-29 MB 2-bit foundation model and Python toolkit for on-device tool calling, structured JSON extraction and text embeddings. It ships inference, LoRA fine-tuning and per-platform builds for phones, wearables, robots, cars and microcontrollers. | View Project → |
| LMCache | ★ 12K | AI Frameworks | Python | kv-cache · llm-inference · inference-optimization · vllm · prefix-caching · cache-offloading · ttft · throughput · long-context · observability · pd-disaggregation · tiered-storage · cuda · rocm · vendor-neutral · serving-engine | LMCache is a vendor-neutral KV cache management layer for LLM inference. It stores and reuses KV cache across CPU RAM, disk and remote backends, cutting TTFT and boosting throughput for long-context, multi-turn agentic and RAG workloads on engines like vLLM. | View Project → |
| OpenRLHF | ★ 10K | AI Frameworks | Python | RLHF · reinforcement-learning · PPO · GRPO · REINFORCE++ · agentic-RL · distributed-training · Ray · vLLM · DeepSpeed · multi-turn-agents · VLM · async-RL · LoRA · reward-model · fine-tuning | OpenRLHF is a Ray + vLLM distributed RLHF framework with an agent-based execution pipeline, supporting PPO, GRPO, REINFORCE++, DAPO, async RL, and VLM/multi-turn agent training for LLMs up to 70B+ parameters. | View Project → |
| adk-go | ★ 8.8K | AI Frameworks | Go | Go · AI agents · multi-agent systems · agent orchestration · tool integration · agent evaluation · MCP · A2A · cloud-native | Google's code-first Go framework for building, orchestrating, evaluating, and deploying AI agents. It supports multi-agent workflows and tool integrations, is optimized for Gemini, and can use other model providers. | View Project → |
| mcp-agent | ★ 8.6K | AI Frameworks | Python | mcp · ai-agents · llm · python · agent-framework · workflows · durable-execution · temporal · orchestrator · multi-agent · mcp-servers · oauth · observability · fastmcp · human-in-the-loop · sdk | Python framework/SDK for building agents on the Model Context Protocol. Fully implements MCP lifecycle (tools, resources, prompts, OAuth, sampling) and composes Anthropic's effective agent patterns, with optional Temporal-backed durable execution and cloud deployment. | View Project → |
| kimi-k3-in-c | ★ 8.3K | AI Frameworks | C | llm-inference · cpu-inference · c99 · zero-dependencies · quantization · mxfp4 · mixture-of-experts · moe · avx2 · simd · memory-efficient · linear-attention · kv-cache · streaming-weights · from-scratch · safetensors | A portable C99 inference engine that runs the 2.78T-parameter Kimi K3 MoE model on one CPU with 8.24 GB peak RAM, streaming a 1.56 TB checkpoint from disk with no BLAS, framework, GPU or runtime dependencies. | View Project → |
| kev | ★ 7.8K | AI Frameworks | Python | decision-model · calibration · lora · qwen3 · inference-engine · serving · evals · benchmarks · probabilities · block-causal-mask · typesafe · fastapi · pytorch · playground · chess · text-classification | Kev trains LoRA adapters plus a pointer readout head on Qwen3 (0.6B/4B/8B) to answer typed yes/no, choice, and score questions over a document in one causal prefill pass, returning calibrated probabilities rather than text. Ships a FastAPI /v1/systemone server, frozen eval suites, and a Next.js playground. | View Project → |
| mlx-lm | ★ 7.2K | AI Frameworks | Python | llm-inference · apple-silicon · mlx · quantization · fine-tuning · lora · gguf · text-generation · prompt-caching · local-llm · openai-compatible · cli · chat-repl · model-conversion · distributed-inference · benchmarking | MLX LM is a Python package for running and fine-tuning LLMs on Apple silicon via MLX. It offers CLI and Python APIs for generation, chat, LoRA/full fine-tuning, quantization, GGUF conversion, prompt caching, and an OpenAI-compatible server. | View Project → |
| DeepSpec | ★ 7.2K | AI Frameworks | Python | speculative-decoding · llm-inference · inference-optimization · draft-model · model-training · eagle3 · dflash · dspark · pytorch · triton · distributed-training · llm-evaluation · throughput · latency · transformers · benchmarking | DeepSpec is a full-stack Python codebase for training and evaluating speculative-decoding draft models (DSpark, DFlash, Eagle3) against targets like Qwen3 and Gemma. It covers data prep, 8-GPU training, and benchmark evaluation, plus released checkpoints. | View Project → |
| shimmy | ★ 5.9K | AI Frameworks | Rust | LLM inference · local AI · inference server · WebGPU · GGUF · OpenAI-compatible API · Rust · GPU acceleration · streaming | Shimmy is a Rust inference server for local GGUF language models, with WebGPU acceleration through Airframe and OpenAI-compatible chat, completion, and streaming APIs. It runs as a single binary without Python or llama.cpp. | View Project → |
| laya-mlx | ★ 5.8K | AI Frameworks | Python | mlx · apple-silicon · inference-runtime · local-inference · on-device-ai · typed-decisions · modernbert · python · fp16 · structured-output · batch-inference · model-conversion · huggingface-hub · macos · low-latency · no-cloud | Native Apple Silicon MLX inference runtime for Laya typed-decision models. Returns choice probabilities, rubric scores and P(true) locally in ~7-14 ms per short question, with no PyTorch, tokenizer decoding or cloud API at runtime. | View Project → |
| higgsfield | ★ 5.8K | AI Frameworks | Jupyter Notebook | distributed-training · gpu-orchestration · llm-training · pytorch · deepspeed · zero-3 · fsdp · fault-tolerance · cluster-management · mlops · multi-node · sharding · llama · training-framework · github-actions · checkpointing | Higgsfield is an open-source GPU orchestration and ML training framework for multi-node training of billion-to-trillion parameter LLMs. It wraps PyTorch FSDP and DeepSpeed ZeRO-3 with node allocation, experiment queuing, monitoring, and GitHub Actions-driven deployment across cloud nodes. | View Project → |
| lemonade | ★ 5.8K | AI Frameworks | C++ | local-llm · llm-inference · openai-api · onnxruntime · llamacpp · npu · rocm · vulkan · multimodal · gguf · text-to-speech · image-generation · mcp-server · amd · local-server · quantization | Lemonade is a local AI server that runs optimized LLMs, speech, and image models on your own GPU/NPU, exposing OpenAI, Anthropic, and Ollama compatible APIs. Ships a CLI, model manager, and MCP server for connecting desktop apps and coding agents to private on-device inference. | View Project → |
| ruoyi-ai | ★ 5.7K | AI Frameworks | Java | multi-agent · rag · mcp · workflow-orchestration · knowledge-base · llm-gateway · vector-database · supervisor-mode · visual-workflow · enterprise · self-hosted · java · spring-boot · langchain4j · docker · skills | Java/Spring Boot enterprise AI platform on Langchain4j: multi-provider LLM management, local RAG with Milvus/Weaviate/Qdrant, MCP tool and Skill integration, visual workflow orchestration, and Supervisor-mode multi-agent coordination with admin and user frontends. | View Project → |
| ruby_llm | ★ 4.4K | AI Frameworks | Ruby | ruby · rails · llm-framework · agents · tool-calling · multimodal · embeddings · rag · structured-output · streaming · image-generation · audio · video · provider-agnostic · mcp · async | RubyLLM is a Ruby-native AI framework giving one consistent API for 18+ providers: chat, streaming, embeddings, RAG, tools, agents, structured output, images, audio, video, and OCR, with first-class Rails integration and cost tracking. | View Project → |
| fast-agent | ★ 3.9K | AI Frameworks | Python | agents · mcp · mcp-client · mcp-server · a2a · acp · agent-skills · cli · tui · multi-agent · workflows · structured-outputs · multimodal · evaluation · ollama · python | Python framework and CLI for building, running and evaluating LLM agents and workflows, with first-class MCP (client/server, sampling, elicitations), Agent Skills, ACP and A2A support. Includes a TUI coding agent plus declarative agent/workflow definitions and broad model provider coverage. | View Project → |
| LazyLLM | ★ 3.9K | AI Frameworks | Python | multi-agent · llm · low-code · rag · fine-tuning · model-deployment · workflow-orchestration · pipeline · agents · tool-calling · vector-database · multimodal · inference · cli | LazyLLM is a Python low-code framework for assembling multi-agent LLM apps from modular pipelines (pipeline, parallel, switch, loop). It bundles RAG, tool-calling agents, one-click deployment, and unified online/local model fine-tuning and inference via vLLM, LightLLM or LMDeploy. | View Project → |
| stable-audio-tools | ★ 3.9K | AI Frameworks | Python | audio-generation · text-to-audio · diffusion-models · latent-diffusion · generative-audio · training · inference · autoencoder · fine-tuning · lora · audio-codec · music-generation · conditioning · gradio · deep-learning · checkpoints | Stability AI's toolkit for training and running generative audio models: latent diffusion, autoencoders, and LMs with text/audio conditioning. Includes training scripts, fine-tuning/LoRA support, and a Gradio UI for Stable Audio Open. | View Project → |
| guppylm | ★ 3.8K | AI Frameworks | Python | llm · tiny-llm · transformer · from-scratch · training · tokenizer · bpe · onnx · wasm · education · pytorch · synthetic-data · inference · colab · quantization | GuppyLM is a ~9M parameter vanilla transformer trained from scratch to chat as a fish persona. It includes data generation, BPE tokenizer training, training loop, inference, ONNX/WASM browser demo, and Colab notebooks, making it a compact reference for building your own tiny LLM. | View Project → |
| sie | ★ 3.3K | AI Frameworks | Python | inference server · model serving · agent infrastructure · embeddings · reranking · semantic search · RAG · OpenAI-compatible API · self-hosted · Kubernetes · multimodal | Self-hosted inference server and production cluster for agent workloads. Serves embedding, retrieval, reranking, generation, OCR, extraction, and multimodal models through an OpenAI-compatible API, with on-demand model loading and integrations for popular AI frameworks and vector stores. | View Project → |
| neo | ★ 3.3K | AI Frameworks | JavaScript | javascript · frontend-framework · multi-threaded · web-workers · zero-build · es-modules · json-first-ui · vdom · agent-memory · graph-rag · mcp-server · multi-agent · self-healing · knowledge-graph · nodejs | Neo.mjs is a multi-threaded JavaScript application engine (worker-based runtime, JSON-first UI, zero-build ES modules), framed as the runtime 'body' hosting an AI agent swarm; the Agent OS, MCP servers and GraphRAG live mostly in the sibling neo-agent-brain repo. | View Project → |
| genaiscript | ★ 2.9K | AI Frameworks | TypeScript | prompt-as-code · llm-orchestration · agents · mcp · rag · vector-search · tool-calling · evals · vscode-extension · cli · javascript · typescript · code-interpreter · content-safety · promptfoo · deprecated | Microsoft GenAIScript is a JavaScript/TypeScript framework for writing LLM prompts as code, orchestrating models, tools, MCP servers and agents. It includes built-in RAG vector search, structured output schemas, evals, and a VS Code extension plus CLI. Note: the repository is marked DEPRECATED. | View Project → |
| AIHOT | ★ 2.5K | AI Frameworks | TypeScript | AI news aggregation · LLM content selection · news clustering · automated briefings · RSS ingestion · industry news · embeddings · MCP · self-hosted | A self-hosted framework for building industry news sites. It collects from configurable sources, uses LLMs to filter and score articles, writes summaries, clusters related coverage into events, and publishes ranked topics and briefings. | View Project → |
| MTPLX | ★ 2.5K | AI Frameworks | Python | apple-silicon · mlx · metal · speculative-decoding · mtp · inference-engine · local-llm · openai-compatible · anthropic-compatible · quantization · macos · qwen · local-server · cli · desktop-app · turbo-decoding | MTPLX is an Apple Silicon LLM inference engine and Mac app that uses Qwen's native multi-token prediction heads for exact speculative decoding (1.6x-2.24x faster decode at any temperature). It serves local models over OpenAI- and Anthropic-compatible APIs for coding agents and chat. | View Project → |
| NanoJev | ★ 2K | AI Frameworks | Python | decision-model · parallel-decoding · game-agents · ViZDoom · reinforcement-learning · SFT · probability-heads · training-pipeline · inference-server · Qwen3 · small-language-model · benchmarks · maze · snake · action-selection · replay-viewer | NanoJev is a 0.6B Qwen3-based replica of Jev: a parallel decision model that returns probability distributions over supplied candidate actions with zero output-token decoding. Ships SFT training configs, an 18.7K-question dataset, an HTTP decision service and ViZDoom/Maze/Snake benchmarks. | View Project → |
| yomo | ★ 1.9K | AI Frameworks | Rust | AI agents · function calling · serverless tools · LLM tool routing · geo-distributed inference · QUIC · streaming · agent skills | Rust framework for building AI agents with function calling and serverless LLM tools. It provides QUIC-based tool routing, an OpenAI-compatible API, and infrastructure for geo-distributed inference. | View Project → |
| LLPhant | ★ 1.7K | AI Frameworks | PHP | generative AI · LLM · PHP · agents · embeddings · RAG · vector databases · question answering · classification · OpenAI-compatible | A PHP framework for building generative AI applications, with support for multiple LLM providers, embeddings, vector stores, agents, and question answering. Integrates with Laravel and Symfony. | View Project → |
| Qwen-Image-2.1 | ★ 1.6K | AI Frameworks | Python | text-to-image · image-editing · diffusion · DiT · RGBA-transparency · 7B-parameters · prompt-rewriting · 2K-resolution · multi-reference · FP8-quantization · ComfyUI · Diffusers · vLLM · SGLang · subject-extraction · inference-optimization | Qwen's 7B text-to-image and image-editing diffusion model (32-layer single-stream DiT) with native RGBA transparency, up to 10 reference images, mask/local edits, and 2K output. Ships Diffusers, ComfyUI, vLLM-Omni and SGLang integrations plus prompt-rewriting checkpoints. | View Project → |
| vllm-mlx | ★ 1.6K | AI Frameworks | Python | llm-inference · mlx · apple-silicon · openai-compatible · anthropic-api · continuous-batching · kv-cache · multimodal · mcp · tool-calling · text-to-speech · speech-to-text · embeddings · reranking · claude-code · local-llm | vLLM-style inference server for Apple Silicon built on MLX, exposing OpenAI /v1/* and Anthropic /v1/messages from one process. Adds continuous batching, paged/prefix KV cache, structured output, MCP tool calling, and multimodal text, vision, audio, embeddings and rerank support. | View Project → |
| npcpy | ★ 1.5K | AI Frameworks | Python | LLM · multimodal · agents · multi-agent · tool calling · MCP · knowledge graphs · model providers | Python framework for building LLM applications with agent, tool-use, and multi-agent orchestration primitives. Supports local and cloud model providers, multimodal workflows, MCP, and knowledge graph pipelines. | View Project → |
| connectonion | ★ 1.5K | AI Frameworks | Python | AI agents · multi-agent · agent framework · agent CLI · tool calling · agent skills · browser automation · agent debugging · agent deployment · approval workflows · persistent memory | Python framework and CLI harness for building, debugging, and deploying tool-using AI agents. Includes browser, shell, email, and file integrations, reusable skills, approval controls, and support for hosting agents that other agents can call. | View Project → |
| xiaozhi-esp32-server-java | ★ 1.4K | AI Frameworks | Java | esp32 · voice-assistant · stt · tts · mcp · spring-ai · websocket · mqtt · rag · iot · smart-home · ota · function-calling · voice-cloning · device-management · knowledge-base | Java enterprise server plus Vue admin console for Xiaozhi ESP32 voice hardware. Multi-LLM (OpenAI/ZhiPu/Ollama/Dify/Coze), local and cloud STT/TTS with voice cloning, WebSocket/MQTT realtime audio, MCP tools, RAG, OTA and device monitoring. | View Project → |
| Gym | ★ 1.2K | AI Frameworks | Python | agent-evaluation · LLM-evaluation · reinforcement-learning · benchmarks · training-environments · rollouts · tool-calling · verifiers · agent-training · model-training · scalable-evaluation | Python library for building environments to evaluate and train LLMs and agents. It provides benchmark environments, agent harnesses, verifiers, rollout collection, and scalable execution for evaluation and reinforcement-learning workflows. | View Project → |
| jevlike | ★ 1.2K | AI Frameworks | Python | one-pass scorer · system-one model · text ranking · attention pooling · training toolkit · evaluation metrics · calibration · byte encoder · frozen encoder · PyTorch · CLI · research starter · vision scorer | Python starter for training a small one-pass scorer that turns text plus a changing list of options into one probability per option, an independent alternative to TypeSafe's Jev. Ships train/eval/predict CLIs, byte and frozen Hugging Face encoder paths, and Doom/chess vision-scoring examples. | View Project → |
| LangChain | ★ 1.1K | AI Frameworks | C# | langchain · csharp · dotnet · llm · rag · embeddings · vector-database · chains · prompt-templates · openai · sdk · nuget · document-loaders · text-splitters | C#/.NET port of LangChain offering composable chains, prompt templates, document loaders, embeddings and vector stores for building LLM and RAG applications, closely mirroring the original Python abstractions. | View Project → |
| CLM | ★ 915 | AI Frameworks | Python | contrastive learning · state-action scoring · decision-making · model serving · verifier · tool selection · candidate ranking · embeddings · fine-tuning | CLM serves a contrastively trained model that scores candidate actions against a state, with an API for typed decisions and ranking. Use it for agent action selection, tool routing, or verifying and ranking generated solutions. | View Project → |
| lmstudio-python | ★ 875 | AI Frameworks | Python | python-sdk · llm · local-llm · lm-studio · inference · chat · tool-use · structured-output · async · websocket · json-schema · speculative-decoding · model-management · developer-tooling | Official Python SDK for LM Studio: connect to a local LM Studio server to run chat, completion, tool-use, and schema-constrained inference with locally hosted LLMs, with sync and async clients plus model load/unload management. | View Project → |
| pipelex | ★ 872 | AI Frameworks | Python | AI workflows · DSL · LLM orchestration · typed pipelines · structured outputs · model routing · document extraction · composable methods | Python framework for declaring and running typed, composable AI methods in .mthds files. It orchestrates multi-step LLM pipelines, model routing, document extraction, and structured outputs. | View Project → |
| swiftide | ★ 786 | AI Frameworks | Rust | Rust · LLM agents · RAG · streaming pipelines · task graphs · tool calling · MCP · vector search · indexing · human-in-the-loop | A Rust framework for building LLM agents, typed task graphs, and streaming RAG pipelines. It includes indexing and query components, tool and MCP integrations, and connectors for LLM providers and vector stores. | View Project → |
| Deuz-SDK | ★ 697 | AI Frameworks | TypeScript | typescript · agent-framework · durable-execution · multi-agent · memory · hybrid-rag · mcp · tool-calling · human-in-the-loop · streaming · guardrails · zero-dependency · edge-runtime · serverless · observability · llm | Zero-runtime-dependency TypeScript SDK for production AI agents: durable resumable runs, long-term memory, hybrid RAG, MCP tool calling, approvals and swarm orchestration over one streaming API for Claude, GPT, Gemini, Grok, Mistral and DeepSeek on Node, Bun, Deno and edge. | View Project → |
| LLMTornado | ★ 641 | AI Frameworks | C# | dotnet · csharp · llm · ai-agents · multi-agent · orchestration · mcp · a2a · rag · vector-database · sdk · multimodal · embeddings · provider-agnostic · local-inference · workflow | .NET SDK for building AI agents and workflows with 30+ provider connectors, MCP and A2A support, vector DB integrations, multimodal IO, and a graph-based agent orchestration API. Works with local runtimes like vLLM, Ollama and LocalAI. | View Project → |
| Swarm | ★ 581 | AI Frameworks | Swift | swift · agents · multi-agent · on-device-inference · foundation-models · mcp · memory · guardrails · streaming · workflows · type-safe-tools · ios · macos · linux · observability · offline-first | Swift-native agent runtime for building AI agents with type-safe @Tool macros, Apple Foundation Models on-device inference, composable sequential/parallel/routed workflows, memory, guardrails, streaming, MCP bridging, and OpenTelemetry tracing. Ships as a Swift Package for iOS, macOS, and Linux. | View Project → |
| Online-RLHF | ★ 546 | AI Frameworks | Python | rlhf · dpo · llm-alignment · online-rlhf · iterative-dpo · reward-modeling · sft · preference-optimization · llama3 · vllm · deepspeed · axolotl · training-pipeline · fine-tuning | Recipe and scripts for online iterative RLHF: SFT, reward modeling, vLLM response generation, reward annotation, and iterative DPO training loops. Reproduces LLaMA3-8B alignment comparable to Llama3-8B-Instruct using only open-source data. | View Project → |
| smg | ★ 543 | AI Frameworks | Rust | llm-gateway · rust · inference-routing · load-balancing · kv-cache-aware · openai-compatible · anthropic-api · grpc · vllm · sglang · tensorrt-llm · mcp · wasm-plugins · multi-tenant · observability · embeddings | SMG is a Rust LLM gateway that unifies OpenAI/Anthropic/Gemini and self-hosted engines (vLLM, SGLang, TensorRT-LLM, TokenSpeed, MLX) behind one API. It adds KV-cache-aware routing, gRPC pipelines, MCP tooling, multi-tenancy and observability for large-scale inference deployments. | View Project → |
| generative-ai-cdk-constructs | ★ 542 | AI Frameworks | TypeScript | aws-cdk · infrastructure-as-code · amazon-bedrock · generative-ai · rag · knowledge-base · sagemaker · opensearch · agents · langchain · step-functions · vector-index · content-generation · summarization · question-answering · typescript | AWS CDK construct library providing multi-service, well-architected patterns for generative AI on AWS: Bedrock, SageMaker model deployment, RAG knowledge bases, OpenSearch vector stores, agents, and batch inference. Use it to define repeatable GenAI infrastructure in TypeScript, Python, Java, Go, or C#. Experimental, n | View Project → |
| ome | ★ 513 | AI Frameworks | Go | kubernetes · llm-serving · model-serving · gpu-scheduling · inference · kubernetes-operator · multi-node · autoscaling · prefill-decode-disaggregation · lora · model-lifecycle · helm · go · benchmarking · mcp-gateway · gpu | OME is a Kubernetes operator for LLM serving: model-as-CRDs, GPU bin-packing scheduling, and automatic runtime selection across SGLang, vLLM, TensorRT-LLM and Triton. Adds prefill-decode disaggregation, multi-node inference, LoRA serving, autoscaling and benchmarking. | View Project → |
| chatluna | ★ 440 | AI Frameworks | TypeScript | chatbot · koishi · llm · langchain · multi-model · mcp-client · agent · plugin · presets · long-term-memory · tts · image-render · web-search · content-moderation · i18n | Koishi chatbot plugin that adds multi-model LLM chat (OpenAI, Claude, Gemini, DeepSeek, Qwen, Ollama and more) via a LangChain-based adapter layer. Offers chat/browse/agent modes, YAML persona presets, MCP client tools, long-term memory, web search and text, voice or image output. | View Project → |
| ai4j | ★ 432 | AI Frameworks | HTML | java · jdk8 · agent-sdk · llm-abstraction · tool-calling · mcp · a2a · rag · coding-agent · cli · tui · spring-boot · vector-store · multi-provider · function-calling · ollama | ai4j is a JDK 8+ Java agentic SDK giving one API across OpenAI, Anthropic, DashScope, DeepSeek, Ollama and more, plus Tool Calling, MCP, A2A, RAG, Agent Runtime and a built-in Coding Agent CLI/TUI/ACP. | View Project → |
| docs | ★ 420 | AI Frameworks | MDX | documentation · langchain · langgraph · langsmith · mintlify · mdx · docs-pipeline · deep-agents · vale · mkdocs-migration | Docs build pipeline and MDX source for docs.langchain.com, covering LangChain, LangGraph, LangSmith, and Deep Agents. It is a documentation monorepo, not an AI library or agent you install. | View Project → |
| awesome-llm-pretraining | ★ 409 | AI Frameworks | — | awesome-list · llm-pretraining · training · datasets · data-curation · moe · model-architecture · scaling · training-strategies · technical-reports · pretraining-data · resource-collection · llm · research | Curated awesome list of LLM pre-training resources: technical reports from Llama, Qwen, DeepSeek and others, training frameworks (Megatron-LM, DeepEP, DeepGEMM), open datasets and data-filtering methods. Useful as a reading and reference index for anyone pre-training or studying LLM training pipelines. | View Project → |
| langchain-google | ★ 404 | AI Frameworks | Python | langchain · google · gemini · vertex-ai · llm · python · integrations · genai · chat-models · embeddings · vector-store · google-cloud · provider-package · rag | Official LangChain integration packages for Google AI: langchain-google-genai (Gemini API), langchain-google-vertexai (Vertex AI), and langchain-google-community. Provides Google chat models, embeddings, vector stores, and tools usable in any LangChain app. | View Project → |
| langchain-aws | ★ 350 | AI Frameworks | Python | langchain · langgraph · aws · amazon-bedrock · bedrock-agentcore · vector-stores · retrievers · rag · agents · checkpointing · memory · sagemaker · kendra · dynamodb · elasticache · python | Monorepo of LangChain and LangGraph integrations for AWS: Bedrock/SageMaker LLMs, AWS vector stores and retrievers for RAG, Bedrock Agents and AgentCore tools, plus DynamoDB/Valkey checkpointers and memory stores. Successor to the AWS components in langchain-community. | View Project → |
| ai | ★ 350 | AI Frameworks | PHP | wordpress · php · wordpress-plugin · abilities-api · gutenberg · block-editor · ai-connectors · content-generation · image-generation · alt-text · comment-moderation · translation · experiment-framework · openai · anthropic · mcp | Official WordPress AI plugin: a modular, opt-in framework that adds AI features (alt text, summarization, translation, image generation, comment moderation) to the Block Editor via the PHP AI Client and Abilities API, with connector plugins for OpenAI, Anthropic and Google. | View Project → |
| b4run | ★ 311 | AI Frameworks | TypeScript | agents · langgraph · langchain · typescript · file-system-routing · typegen · workflows · tools · durable-threads · sandbox-execution · approval-flows · dev-server · build-artifacts · nodejs | B4.run is a TypeScript meta-framework that wraps LangGraph.js with file-system routes, generated route/state/tool types, workspace sandboxing, approval gates, durable threads, and fixture-backed tests to emit runnable Node servers and Dockerfiles. | View Project → |
| ContinualLM | ★ 295 | AI Frameworks | Python | continual-learning · catastrophic-forgetting · domain-adaptive-pretraining · language-models · knowledge-transfer · transfer-learning · transformer · soft-masking · adapters · prompts · knowledge-distillation · fine-tuning · pytorch · research · nlp | PyTorch framework for continual learning of language models: implements DAS, CPT, DGA, EWC, HAT, DER++ and baselines for sequential domain-adaptive pretraining with end-task fine-tuning, forgetting-rate tools, and Hugging Face checkpoints. | View Project → |
| llama_ros | ★ 264 | AI Frameworks | C++ | ros2 · llama.cpp · gguf · llm · vlm · robotics · local-inference · on-device · embeddings · rerank · langchain · behavior-trees · multimodal · llava · lora · cuda | ROS 2 packages that wrap llama.cpp and llava.cpp, exposing GGUF LLMs and VLMs as ROS 2 nodes with launch files, behavior-tree nodes, LangChain/RAG integration and LoRA/grammar support for local robotics inference. | View Project → |
| wavefront | ★ 200 | AI Frameworks | Python | AI middleware · AI workflows · agent orchestration · RAG · MCP connectors · guardrails · observability · enterprise AI · data integrations · RBAC · voice agents | Wavefront is an open-source middleware platform for building and operating enterprise AI agents, workflows, and RAG applications. It provides data integrations, MCP connectors, access controls, and observability. | View Project → |
| agent-kernel | ★ 191 | AI Frameworks | Python | ai-agents · multi-agent · agent-orchestration · mcp · a2a · ag-ui · enterprise · guardrails · observability · session-management · serverless · kubernetes · terraform · multi-cloud · rag · skills | Agent Kernel is a Python platform layer for running, orchestrating and deploying production AI agents. It runs OpenAI Agents SDK, LangGraph, CrewAI and Google ADK side by side, adds guardrails, sessions, RAG, sandboxing, channels and MCP/A2A/AG-UI, and deploys to AWS, Azure, GCP or Kubernetes via Terraform and Helm. | View Project → |
| yoagent | ★ 179 | AI Frameworks | Rust | Rust · agent loop · tool calling · streaming · coding agents · multi-agent · MCP · OpenAPI · local models · session branching · middleware · LLM providers | A Rust framework for building tool-using LLM agents, with streaming support across seven protocols, built-in tools, MCP and OpenAPI integrations, sub-agents, and session management. Includes a terminal coding-agent example. | View Project → |