Vibe Coding Discover

Category

AI Frameworks

Core frameworks for building LLM applications.

108 projects

tensorflow

★ 200K

An end-to-end machine-learning framework with Python and C++ APIs for building, training, and deploying models across CPUs, GPUs, and other devices.

AI Frameworks | C++ · machine-learning · deep-learning

View Project →

ollama

★ 181K

Ollama is a Go-based local LLM runtime that downloads, runs, and serves open models (Gemma, Qwen, DeepSeek, Llama) via a CLI and REST API on port 11434. It powers local inference for coding agents, chat UIs, and RAG apps, and supports custom Modelfiles and imports.

AI Frameworks | Go · llm · local-llm

View Project →

transformers

★ 166K

Hugging Face Transformers is the model-definition framework for state-of-the-art text, vision, audio, video and multimodal models. It offers a unified Pipeline/Trainer API over 1M+ Hub checkpoints and feeds model definitions to vLLM, llama.cpp and training stacks.

AI Frameworks | Python · transformers · model-definition

View Project →

dify

★ 157K

Dify is an open-source LLM app platform combining a visual agentic workflow builder, RAG pipeline, prompt IDE, agent runtime, and LLMOps. It supports hundreds of models, MCP servers, and tools, and self-hosts via Docker with APIs for integration.

AI Frameworks | TypeScript · llm · agents

View Project →

langflow

★ 155K

Langflow is an open-source visual platform for building and deploying LLM agents and workflows. It combines a drag-and-drop React Flow canvas with Python component customization, then ships each flow as a REST API, an MCP server, or exportable JSON.

AI Frameworks | Python · visual-builder · low-code

View Project →

langchain

★ 147K

LangChain is the Python/JS framework for building LLM apps and agents: a standard interface for chat models, embeddings, vector stores, tools, and retrievers, plus chains, structured output, and streaming. Paired with LangGraph for controllable agent workflows.

AI Frameworks | Python · agents · llm

View Project →

llama.cpp

★ 129K

llama.cpp is a dependency-free C/C++ LLM and VLM inference engine built on ggml. It runs quantized GGUF models on CPU, CUDA, Metal, Vulkan and many other backends, and ships CLI, server, web UI and OpenAI-compatible API tools.

AI Frameworks | C++ · llm-inference · gguf

View Project →

llama.cpp

★ 129K

High-performance C/C++ inference for local LLMs across CPUs, GPUs, and devices.

AI Frameworks | C++ · llama.cpp · Inference

View Project →

vllm

★ 92K

vLLM is a high-throughput, memory-efficient LLM inference and serving engine built on PagedAttention, continuous batching, and optimized CUDA/ROCm kernels. It exposes an OpenAI-compatible server with quantization, LoRA, speculative decoding, and disaggregated serving across 200+ model architectures.

AI Frameworks | Python · inference-engine · llm-serving

View Project →

unsloth

★ 77K

Unsloth is a Python training and inference stack plus desktop/web UI for running, fine-tuning and RL-training LLMs and diffusion models 2x faster with less VRAM. Supports LoRA/QLoRA/full fine-tuning, GGUF/MLX export, local OpenAI-compatible serving, RAG and agent hooks. Apache-2.0.

AI Frameworks | Python · fine-tuning · llm

View Project →

annotated_deep_learning_paper_implementations

★ 67K

LabML's annotated deep learning library: 60+ readable PyTorch implementations of papers with side-by-side notes, covering transformers (GPT, ViT, XL), LoRA, diffusion/Stable Diffusion, GANs, optimizers, and RL. Best as a learning and reference resource rather than a production framework.

AI Frameworks | Python · transformers · pytorch

View Project →

nanoGPT

★ 63K

Minimal, readable ~300-line PyTorch repo for training and finetuning GPT-2-scale models; reproduces GPT-2 124M on OpenWebText. Author marks it deprecated in favor of the newer nanochat, but it remains a compact hacking/education base.

AI Frameworks | Python · gpt · llm-training

View Project →

litellm

★ 59K

LiteLLM is an open-source AI gateway and Python SDK that calls 100+ LLM providers (OpenAI, Anthropic, Bedrock, Vertex, vLLM, Ollama) in OpenAI format. The proxy adds virtual keys, spend tracking, guardrails, load balancing, logging, plus MCP tool and A2A agent routing.

AI Frameworks | Python · ai-gateway · llm-gateway

View Project →

llama_index

★ 52K

LlamaIndex is a Python framework for building LLM apps: data connectors, indices, retrievers and query engines for RAG, plus agent/workflow abstractions. 300+ integrations let you swap LLMs, embeddings and vector stores; the team now also pushes the hosted LlamaParse document platform.

AI Frameworks | Python · rag · llm

View Project →

LocalAI

★ 49K

LocalAI is a self-hosted, OpenAI/Anthropic-compatible inference engine that runs LLMs, vision, voice, image and video models on CPU or GPU. Backends like llama.cpp, vLLM and whisper.cpp are pulled on demand, and it also ships agents, RAG and MCP support.

AI Frameworks | Go · llm-inference · local-inference

View Project →

agno

★ 42K

Agno is a Python framework plus AgentOS runtime for building, serving, and managing agent platforms. It adds REST/SSE endpoints, storage, memory, RAG knowledge, 100+ toolkits, human approval, scheduling, RBAC and OpenTelemetry observability, with Docker and cloud deploy templates.

AI Frameworks | Python · agents · multi-agent

View Project →

langgraph

★ 42K

LangGraph is a low-level Python orchestration framework for building stateful, long-running agents as graphs. It adds durable execution, checkpointing, human-in-the-loop interrupts, and memory, with optional LangSmith tracing and deployment.

AI Frameworks | Python · agents · multi-agent

View Project →

dspy

★ 38K

DSPy is a Python framework for programming rather than prompting LLMs: you write declarative modules and signatures, then compilers/optimizers (e.g. GEPA, MIPRO) tune prompts and weights against your metric. Useful for building and optimizing RAG pipelines, classifiers, and agent loops.

AI Frameworks | Python · dspy · prompt-optimization

View Project →

colibri

★ 38K

Colibrì is a pure-C, zero-dependency inference engine that runs frontier MoE models (GLM-5.x 744B, Kimi K3 2.8T, DeepSeek V4, Qwen3.x) on consumer hardware by streaming disk-resident experts into a unified VRAM/RAM/NVMe tier hierarchy. It offers chat/serve/web front ends plus CPU, CUDA, Metal and Vulkan backends.

AI Frameworks | C · inference-engine · moe

View Project →

CopilotKit

★ 37K

CopilotKit is a TypeScript SDK for building agent-native apps: chat UI, generative UI, shared state, and human-in-the-loop across React, Angular, Vue, React Native, Slack, and Teams. It also maintains the AG-UI protocol connecting agent frameworks to frontends.

AI Frameworks | TypeScript · agent-native · generative-ui

View Project →

sglang

★ 36K

SGLang is a high-performance serving framework for LLMs and multimodal models, offering RadixAttention prefix caching, prefill-decode disaggregation, speculative decoding, and quantization over an OpenAI-compatible API across NVIDIA, AMD, TPU, and CPU hardware.

AI Frameworks | Python · llm-serving · inference-engine

View Project →

airllm

★ 35K

AirLLM is a Python inference engine that runs very large LLMs (70B up to 671B/2.8T) on a single 4GB GPU without quantization, distillation, or pruning by streaming model layers from disk. Optional 4/8-bit block-wise compression gives ~3x speedup, and it supports macOS/Apple Silicon plus low-VRAM training of 125B models

AI Frameworks | Jupyter Notebook · llm-inference · memory-optimization

View Project →

VoiceStudio

★ 35K

Fully local, open-source ElevenLabs alternative for voice cloning, voice design, dubbing, dictation, transcription and audiobooks in 646 languages. Ships a desktop app (Electron/Tauri), model manager, local API and an MCP server for agent integration.

AI Frameworks | Python · voice-cloning · text-to-speech

View Project →

agentscope

★ 33K

AgentScope 2.0 is Alibaba Tongyi's Python framework for building and running production agents: ReAct agents, toolkits over Python/MCP/skills, multi-agent teams, realtime voice, sandboxed execution, HITL, memory, plus a FastAPI agent service with web UI and IM channels.

AI Frameworks | Python · multi-agent · agent-framework

View Project →

modular

★ 30K

Modular's open-source platform hosting MAX (an inference framework with an OpenAI-compatible server and accelerator kernels) and the Mojo language, including its compiler, standard library, and Python model pipelines.

AI Frameworks | Mojo · max · llm-inference

View Project →

semantic-kernel

★ 29K

Microsoft's model-agnostic SDK for building AI agents and multi-agent systems in Python, .NET, and Java. Ships plugins/tool calling, memory and vector connectors, planners, and MCP support; now succeeded by Microsoft Agent Framework.

AI Frameworks | C# · ai-agents · multi-agent

View Project →

mastra

★ 28K

Mastra is a TypeScript framework for building AI agents and LLM apps: agents, graph workflows, RAG, memory, MCP servers, evals, and observability, with model routing across 40+ providers and React/Next.js/Node integration.

AI Frameworks | TypeScript · ai-agents · workflows

View Project →

ai

★ 27K

Vercel AI SDK is a provider-agnostic TypeScript toolkit for building AI apps and agents: unified model APIs, streaming text, structured output, tool loops, and framework-agnostic UI hooks for React, Next.js, Svelte and Vue.

AI Frameworks | TypeScript · ai-sdk · llm

View Project →

haystack

★ 27K

Haystack is deepset's open-source Python framework for building production LLM apps: modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Model- and vendor-agnostic, with built-in components for RAG, semantic search, evaluation, and async streaming.

AI Frameworks | Python · llm-orchestration · rag

View Project →

laya

★ 26K

Python library for fast, calibrated typed decisions over text, including classification, scoring, and yes/no probabilities in a single model pass. It routes requests to English or multilingual checkpoints and supports HTTP, MCP, and LangChain integrations.

AI Frameworks | Python · decision-model · classification

View Project →

omlx

★ 22K

oMLX is an MLX-based LLM inference server for Apple Silicon with continuous batching, tiered hot-RAM/cold-SSD KV cache, multi-model serving, and VLMs/embeddings/rerankers behind an OpenAI-compatible API, managed from a macOS menu bar app and web dashboard.

AI Frameworks | Python · llm-inference · apple-silicon

View Project →

onnxruntime

★ 22K

Cross-platform ML inference and training accelerator. Runs models from PyTorch, TensorFlow, and classical ML libraries via ONNX, applying graph optimizations and hardware acceleration on CPU, GPU, or NPU. Widely used runtime for deploying fast, low-cost model inference.

AI Frameworks | C++ · onnx · inference-engine

View Project →

peft

★ 22K

Hugging Face PEFT implements parameter-efficient fine-tuning methods (LoRA, QLoRA, IA3, prompt tuning, adapters and more) on top of Transformers, Diffusers, Accelerate and TRL. It lets you train and serve large models with a fraction of the GPU memory and storage, with only small adapter checkpoints.

AI Frameworks | Python · peft · lora

View Project →

pydantic-ai

★ 20K

Pydantic AI is a Python AI agent framework from the Pydantic team: a typed agent loop with validated structured outputs, tools, dependency injection, MCP, multi-agent support, realtime voice, image generation, embeddings and evals, with provider-agnostic model switching.

AI Frameworks | Python · agent-framework · llm

View Project →

json-render

★ 18K

Vercel Labs' Generative UI framework: an LLM emits JSON specs constrained to a Zod-typed component catalog, which json-render streams and renders safely across React, Vue, Svelte, Solid, React Native, Next.js, Remotion, PDF, email, 3D and terminal targets.

AI Frameworks | TypeScript · generative-ui · json-spec

View Project →

rocketride-server

★ 18K

MIT-licensed AI pipeline engine with a multithreaded C++ runtime and 50+ Python-extensible nodes. Compose LLM, RAG and multi-agent workflows as portable JSON in VS Code or via Python/TypeScript/MCP SDKs, backed by 13+ model providers and 9 vector DBs.

AI Frameworks | Python · ai-pipeline · llm-workflows

View Project →

instructor

★ 14K

Structured LLM outputs validated by Pydantic, with retries and streaming.

AI Frameworks | Python · Structured Output · Pydantic

View Project →

agent-framework

★ 14K

Microsoft's open multi-language SDK for building, orchestrating and hosting production AI agents and multi-agent workflows in Python and .NET. Ships graph workflows, middleware, OpenTelemetry observability, declarative YAML agents and Foundry hosting.

AI Frameworks | Python · agents · multi-agent

View Project →

semantica

★ 13K

Semantica is a Python graph-native context and knowledge-graph layer for AI systems: ingest data, build context graphs, and run deterministic graph reasoning with W3C PROV-O provenance. Adds graph RAG, decision audit trails, ontology governance, and MCP/LangChain/CrewAI integrations on top of your existing LLM stack.

AI Frameworks | Python · knowledge-graph · context-graph

View Project →

eino

★ 13K

Eino is ByteDance's Go LLM application framework, modeled on LangChain and Google ADK. It offers component abstractions (ChatModel, Tool, Retriever), a graph/compose orchestration layer with automatic streaming, an agent development kit with ReAct and multi-agent patterns, and interrupt/resume for human-in-the-loop.

AI Frameworks | Go · golang · llm-framework

View Project →

langchain4j

★ 13K

LangChain4j is an idiomatic Java library for building LLM-powered JVM apps, with a unified API over 20+ model providers and 30+ embedding stores. It covers tool calling (incl. MCP), agents, RAG pipelines, chat memory and prompt templating, with Spring Boot and Quarkus integrations.

AI Frameworks | Java · jvm · llm

View Project →

needle

★ 13K

Needle is an 8-29 MB 2-bit foundation model and Python toolkit for on-device tool calling, structured JSON extraction and text embeddings. It ships inference, LoRA fine-tuning and per-platform builds for phones, wearables, robots, cars and microcontrollers.

AI Frameworks | Python · on-device-ai · edge-ai

View Project →

LMCache

★ 12K

LMCache is a vendor-neutral KV cache management layer for LLM inference. It stores and reuses KV cache across CPU RAM, disk and remote backends, cutting TTFT and boosting throughput for long-context, multi-turn agentic and RAG workloads on engines like vLLM.

AI Frameworks | Python · kv-cache · llm-inference

View Project →

OpenRLHF

★ 10K

OpenRLHF is a Ray + vLLM distributed RLHF framework with an agent-based execution pipeline, supporting PPO, GRPO, REINFORCE++, DAPO, async RL, and VLM/multi-turn agent training for LLMs up to 70B+ parameters.

AI Frameworks | Python · RLHF · reinforcement-learning

View Project →

adk-go

★ 8.8K

Google's code-first Go framework for building, orchestrating, evaluating, and deploying AI agents. It supports multi-agent workflows and tool integrations, is optimized for Gemini, and can use other model providers.

AI Frameworks | Go · AI agents · multi-agent systems

View Project →

mcp-agent

★ 8.6K

Python framework/SDK for building agents on the Model Context Protocol. Fully implements MCP lifecycle (tools, resources, prompts, OAuth, sampling) and composes Anthropic's effective agent patterns, with optional Temporal-backed durable execution and cloud deployment.

AI Frameworks | Python · mcp · ai-agents

View Project →

kimi-k3-in-c

★ 8.3K

A portable C99 inference engine that runs the 2.78T-parameter Kimi K3 MoE model on one CPU with 8.24 GB peak RAM, streaming a 1.56 TB checkpoint from disk with no BLAS, framework, GPU or runtime dependencies.

AI Frameworks | C · llm-inference · cpu-inference

View Project →

kev

★ 7.8K

Kev trains LoRA adapters plus a pointer readout head on Qwen3 (0.6B/4B/8B) to answer typed yes/no, choice, and score questions over a document in one causal prefill pass, returning calibrated probabilities rather than text. Ships a FastAPI /v1/systemone server, frozen eval suites, and a Next.js playground.

AI Frameworks | Python · decision-model · calibration

View Project →

mlx-lm

★ 7.2K

MLX LM is a Python package for running and fine-tuning LLMs on Apple silicon via MLX. It offers CLI and Python APIs for generation, chat, LoRA/full fine-tuning, quantization, GGUF conversion, prompt caching, and an OpenAI-compatible server.

AI Frameworks | Python · llm-inference · apple-silicon

View Project →

DeepSpec

★ 7.2K

DeepSpec is a full-stack Python codebase for training and evaluating speculative-decoding draft models (DSpark, DFlash, Eagle3) against targets like Qwen3 and Gemma. It covers data prep, 8-GPU training, and benchmark evaluation, plus released checkpoints.

AI Frameworks | Python · speculative-decoding · llm-inference

View Project →

shimmy

★ 5.9K

Shimmy is a Rust inference server for local GGUF language models, with WebGPU acceleration through Airframe and OpenAI-compatible chat, completion, and streaming APIs. It runs as a single binary without Python or llama.cpp.

AI Frameworks | Rust · LLM inference · local AI

View Project →

laya-mlx

★ 5.8K

Native Apple Silicon MLX inference runtime for Laya typed-decision models. Returns choice probabilities, rubric scores and P(true) locally in ~7-14 ms per short question, with no PyTorch, tokenizer decoding or cloud API at runtime.

AI Frameworks | Python · mlx · apple-silicon

View Project →

higgsfield

★ 5.8K

Higgsfield is an open-source GPU orchestration and ML training framework for multi-node training of billion-to-trillion parameter LLMs. It wraps PyTorch FSDP and DeepSpeed ZeRO-3 with node allocation, experiment queuing, monitoring, and GitHub Actions-driven deployment across cloud nodes.

AI Frameworks | Jupyter Notebook · distributed-training · gpu-orchestration

View Project →

lemonade

★ 5.8K

Lemonade is a local AI server that runs optimized LLMs, speech, and image models on your own GPU/NPU, exposing OpenAI, Anthropic, and Ollama compatible APIs. Ships a CLI, model manager, and MCP server for connecting desktop apps and coding agents to private on-device inference.

AI Frameworks | C++ · local-llm · llm-inference

View Project →

ruoyi-ai

★ 5.7K

Java/Spring Boot enterprise AI platform on Langchain4j: multi-provider LLM management, local RAG with Milvus/Weaviate/Qdrant, MCP tool and Skill integration, visual workflow orchestration, and Supervisor-mode multi-agent coordination with admin and user frontends.

AI Frameworks | Java · multi-agent · rag

View Project →

ruby_llm

★ 4.4K

RubyLLM is a Ruby-native AI framework giving one consistent API for 18+ providers: chat, streaming, embeddings, RAG, tools, agents, structured output, images, audio, video, and OCR, with first-class Rails integration and cost tracking.

AI Frameworks | Ruby · rails · llm-framework

View Project →

fast-agent

★ 3.9K

Python framework and CLI for building, running and evaluating LLM agents and workflows, with first-class MCP (client/server, sampling, elicitations), Agent Skills, ACP and A2A support. Includes a TUI coding agent plus declarative agent/workflow definitions and broad model provider coverage.

AI Frameworks | Python · agents · mcp

View Project →

LazyLLM

★ 3.9K

LazyLLM is a Python low-code framework for assembling multi-agent LLM apps from modular pipelines (pipeline, parallel, switch, loop). It bundles RAG, tool-calling agents, one-click deployment, and unified online/local model fine-tuning and inference via vLLM, LightLLM or LMDeploy.

AI Frameworks | Python · multi-agent · llm

View Project →

stable-audio-tools

★ 3.9K

Stability AI's toolkit for training and running generative audio models: latent diffusion, autoencoders, and LMs with text/audio conditioning. Includes training scripts, fine-tuning/LoRA support, and a Gradio UI for Stable Audio Open.

AI Frameworks | Python · audio-generation · text-to-audio

View Project →

guppylm

★ 3.8K

GuppyLM is a ~9M parameter vanilla transformer trained from scratch to chat as a fish persona. It includes data generation, BPE tokenizer training, training loop, inference, ONNX/WASM browser demo, and Colab notebooks, making it a compact reference for building your own tiny LLM.

AI Frameworks | Python · llm · tiny-llm

View Project →

sie

★ 3.3K

Self-hosted inference server and production cluster for agent workloads. Serves embedding, retrieval, reranking, generation, OCR, extraction, and multimodal models through an OpenAI-compatible API, with on-demand model loading and integrations for popular AI frameworks and vector stores.

AI Frameworks | Python · inference server · model serving

View Project →

neo

★ 3.3K

Neo.mjs is a multi-threaded JavaScript application engine (worker-based runtime, JSON-first UI, zero-build ES modules), framed as the runtime 'body' hosting an AI agent swarm; the Agent OS, MCP servers and GraphRAG live mostly in the sibling neo-agent-brain repo.

AI Frameworks | JavaScript · frontend-framework · multi-threaded

View Project →

genaiscript

★ 2.9K

Microsoft GenAIScript is a JavaScript/TypeScript framework for writing LLM prompts as code, orchestrating models, tools, MCP servers and agents. It includes built-in RAG vector search, structured output schemas, evals, and a VS Code extension plus CLI. Note: the repository is marked DEPRECATED.

AI Frameworks | TypeScript · prompt-as-code · llm-orchestration

View Project →

AIHOT

★ 2.5K

A self-hosted framework for building industry news sites. It collects from configurable sources, uses LLMs to filter and score articles, writes summaries, clusters related coverage into events, and publishes ranked topics and briefings.

AI Frameworks | TypeScript · AI news aggregation · LLM content selection

View Project →

MTPLX

★ 2.5K

MTPLX is an Apple Silicon LLM inference engine and Mac app that uses Qwen's native multi-token prediction heads for exact speculative decoding (1.6x-2.24x faster decode at any temperature). It serves local models over OpenAI- and Anthropic-compatible APIs for coding agents and chat.

AI Frameworks | Python · apple-silicon · mlx

View Project →

NanoJev

★ 2K

NanoJev is a 0.6B Qwen3-based replica of Jev: a parallel decision model that returns probability distributions over supplied candidate actions with zero output-token decoding. Ships SFT training configs, an 18.7K-question dataset, an HTTP decision service and ViZDoom/Maze/Snake benchmarks.

AI Frameworks | Python · decision-model · parallel-decoding

View Project →

yomo

★ 1.9K

Rust framework for building AI agents with function calling and serverless LLM tools. It provides QUIC-based tool routing, an OpenAI-compatible API, and infrastructure for geo-distributed inference.

AI Frameworks | Rust · AI agents · function calling

View Project →

LLPhant

★ 1.7K

A PHP framework for building generative AI applications, with support for multiple LLM providers, embeddings, vector stores, agents, and question answering. Integrates with Laravel and Symfony.

AI Frameworks | PHP · generative AI · LLM

View Project →

Qwen-Image-2.1

★ 1.6K

Qwen's 7B text-to-image and image-editing diffusion model (32-layer single-stream DiT) with native RGBA transparency, up to 10 reference images, mask/local edits, and 2K output. Ships Diffusers, ComfyUI, vLLM-Omni and SGLang integrations plus prompt-rewriting checkpoints.

AI Frameworks | Python · text-to-image · image-editing

View Project →

vllm-mlx

★ 1.6K

vLLM-style inference server for Apple Silicon built on MLX, exposing OpenAI /v1/* and Anthropic /v1/messages from one process. Adds continuous batching, paged/prefix KV cache, structured output, MCP tool calling, and multimodal text, vision, audio, embeddings and rerank support.

AI Frameworks | Python · llm-inference · mlx

View Project →

npcpy

★ 1.5K

Python framework for building LLM applications with agent, tool-use, and multi-agent orchestration primitives. Supports local and cloud model providers, multimodal workflows, MCP, and knowledge graph pipelines.

AI Frameworks | Python · LLM · multimodal

View Project →

connectonion

★ 1.5K

Python framework and CLI harness for building, debugging, and deploying tool-using AI agents. Includes browser, shell, email, and file integrations, reusable skills, approval controls, and support for hosting agents that other agents can call.

AI Frameworks | Python · AI agents · multi-agent

View Project →

xiaozhi-esp32-server-java

★ 1.4K

Java enterprise server plus Vue admin console for Xiaozhi ESP32 voice hardware. Multi-LLM (OpenAI/ZhiPu/Ollama/Dify/Coze), local and cloud STT/TTS with voice cloning, WebSocket/MQTT realtime audio, MCP tools, RAG, OTA and device monitoring.

AI Frameworks | Java · esp32 · voice-assistant

View Project →

Gym

★ 1.2K

Python library for building environments to evaluate and train LLMs and agents. It provides benchmark environments, agent harnesses, verifiers, rollout collection, and scalable execution for evaluation and reinforcement-learning workflows.

AI Frameworks | Python · agent-evaluation · LLM-evaluation

View Project →

jevlike

★ 1.2K

Python starter for training a small one-pass scorer that turns text plus a changing list of options into one probability per option, an independent alternative to TypeSafe's Jev. Ships train/eval/predict CLIs, byte and frozen Hugging Face encoder paths, and Doom/chess vision-scoring examples.

AI Frameworks | Python · one-pass scorer · system-one model

View Project →

LangChain

★ 1.1K

C#/.NET port of LangChain offering composable chains, prompt templates, document loaders, embeddings and vector stores for building LLM and RAG applications, closely mirroring the original Python abstractions.

AI Frameworks | C# · langchain · csharp

View Project →

CLM

★ 915

CLM serves a contrastively trained model that scores candidate actions against a state, with an API for typed decisions and ranking. Use it for agent action selection, tool routing, or verifying and ranking generated solutions.

AI Frameworks | Python · contrastive learning · state-action scoring

View Project →

lmstudio-python

★ 875

Official Python SDK for LM Studio: connect to a local LM Studio server to run chat, completion, tool-use, and schema-constrained inference with locally hosted LLMs, with sync and async clients plus model load/unload management.

AI Frameworks | Python · python-sdk · llm

View Project →

pipelex

★ 872

Python framework for declaring and running typed, composable AI methods in .mthds files. It orchestrates multi-step LLM pipelines, model routing, document extraction, and structured outputs.

AI Frameworks | Python · AI workflows · DSL

View Project →

swiftide

★ 786

A Rust framework for building LLM agents, typed task graphs, and streaming RAG pipelines. It includes indexing and query components, tool and MCP integrations, and connectors for LLM providers and vector stores.

AI Frameworks | Rust · LLM agents · RAG

View Project →

Deuz-SDK

★ 697

Zero-runtime-dependency TypeScript SDK for production AI agents: durable resumable runs, long-term memory, hybrid RAG, MCP tool calling, approvals and swarm orchestration over one streaming API for Claude, GPT, Gemini, Grok, Mistral and DeepSeek on Node, Bun, Deno and edge.

AI Frameworks | TypeScript · agent-framework · durable-execution

View Project →

LLMTornado

★ 641

.NET SDK for building AI agents and workflows with 30+ provider connectors, MCP and A2A support, vector DB integrations, multimodal IO, and a graph-based agent orchestration API. Works with local runtimes like vLLM, Ollama and LocalAI.

AI Frameworks | C# · dotnet · csharp

View Project →

Swarm

★ 581

Swift-native agent runtime for building AI agents with type-safe @Tool macros, Apple Foundation Models on-device inference, composable sequential/parallel/routed workflows, memory, guardrails, streaming, MCP bridging, and OpenTelemetry tracing. Ships as a Swift Package for iOS, macOS, and Linux.

AI Frameworks | Swift · agents · multi-agent

View Project →

Online-RLHF

★ 546

Recipe and scripts for online iterative RLHF: SFT, reward modeling, vLLM response generation, reward annotation, and iterative DPO training loops. Reproduces LLaMA3-8B alignment comparable to Llama3-8B-Instruct using only open-source data.

AI Frameworks | Python · rlhf · dpo

View Project →

smg

★ 543

SMG is a Rust LLM gateway that unifies OpenAI/Anthropic/Gemini and self-hosted engines (vLLM, SGLang, TensorRT-LLM, TokenSpeed, MLX) behind one API. It adds KV-cache-aware routing, gRPC pipelines, MCP tooling, multi-tenancy and observability for large-scale inference deployments.

AI Frameworks | Rust · llm-gateway · inference-routing

View Project →

generative-ai-cdk-constructs

★ 542

AWS CDK construct library providing multi-service, well-architected patterns for generative AI on AWS: Bedrock, SageMaker model deployment, RAG knowledge bases, OpenSearch vector stores, agents, and batch inference. Use it to define repeatable GenAI infrastructure in TypeScript, Python, Java, Go, or C#. Experimental, n

AI Frameworks | TypeScript · aws-cdk · infrastructure-as-code

View Project →

ome

★ 513

OME is a Kubernetes operator for LLM serving: model-as-CRDs, GPU bin-packing scheduling, and automatic runtime selection across SGLang, vLLM, TensorRT-LLM and Triton. Adds prefill-decode disaggregation, multi-node inference, LoRA serving, autoscaling and benchmarking.

AI Frameworks | Go · kubernetes · llm-serving

View Project →

chatluna

★ 440

Koishi chatbot plugin that adds multi-model LLM chat (OpenAI, Claude, Gemini, DeepSeek, Qwen, Ollama and more) via a LangChain-based adapter layer. Offers chat/browse/agent modes, YAML persona presets, MCP client tools, long-term memory, web search and text, voice or image output.

AI Frameworks | TypeScript · chatbot · koishi

View Project →

ai4j

★ 432

ai4j is a JDK 8+ Java agentic SDK giving one API across OpenAI, Anthropic, DashScope, DeepSeek, Ollama and more, plus Tool Calling, MCP, A2A, RAG, Agent Runtime and a built-in Coding Agent CLI/TUI/ACP.

AI Frameworks | HTML · java · jdk8

View Project →

docs

★ 420

Docs build pipeline and MDX source for docs.langchain.com, covering LangChain, LangGraph, LangSmith, and Deep Agents. It is a documentation monorepo, not an AI library or agent you install.

AI Frameworks | MDX · documentation · langchain

View Project →

awesome-llm-pretraining

★ 409

Curated awesome list of LLM pre-training resources: technical reports from Llama, Qwen, DeepSeek and others, training frameworks (Megatron-LM, DeepEP, DeepGEMM), open datasets and data-filtering methods. Useful as a reading and reference index for anyone pre-training or studying LLM training pipelines.

AI Frameworks | awesome-list · llm-pretraining

View Project →

langchain-google

★ 404

Official LangChain integration packages for Google AI: langchain-google-genai (Gemini API), langchain-google-vertexai (Vertex AI), and langchain-google-community. Provides Google chat models, embeddings, vector stores, and tools usable in any LangChain app.

AI Frameworks | Python · langchain · google

View Project →

langchain-aws

★ 350

Monorepo of LangChain and LangGraph integrations for AWS: Bedrock/SageMaker LLMs, AWS vector stores and retrievers for RAG, Bedrock Agents and AgentCore tools, plus DynamoDB/Valkey checkpointers and memory stores. Successor to the AWS components in langchain-community.

AI Frameworks | Python · langchain · langgraph

View Project →

ai

★ 350

Official WordPress AI plugin: a modular, opt-in framework that adds AI features (alt text, summarization, translation, image generation, comment moderation) to the Block Editor via the PHP AI Client and Abilities API, with connector plugins for OpenAI, Anthropic and Google.

AI Frameworks | PHP · wordpress · wordpress-plugin

View Project →

b4run

★ 311

B4.run is a TypeScript meta-framework that wraps LangGraph.js with file-system routes, generated route/state/tool types, workspace sandboxing, approval gates, durable threads, and fixture-backed tests to emit runnable Node servers and Dockerfiles.

AI Frameworks | TypeScript · agents · langgraph

View Project →

ContinualLM

★ 295

PyTorch framework for continual learning of language models: implements DAS, CPT, DGA, EWC, HAT, DER++ and baselines for sequential domain-adaptive pretraining with end-task fine-tuning, forgetting-rate tools, and Hugging Face checkpoints.

AI Frameworks | Python · continual-learning · catastrophic-forgetting

View Project →

llama_ros

★ 264

ROS 2 packages that wrap llama.cpp and llava.cpp, exposing GGUF LLMs and VLMs as ROS 2 nodes with launch files, behavior-tree nodes, LangChain/RAG integration and LoRA/grammar support for local robotics inference.

AI Frameworks | C++ · ros2 · llama.cpp

View Project →

wavefront

★ 200

Wavefront is an open-source middleware platform for building and operating enterprise AI agents, workflows, and RAG applications. It provides data integrations, MCP connectors, access controls, and observability.

AI Frameworks | Python · AI middleware · AI workflows

View Project →

agent-kernel

★ 191

Agent Kernel is a Python platform layer for running, orchestrating and deploying production AI agents. It runs OpenAI Agents SDK, LangGraph, CrewAI and Google ADK side by side, adds guardrails, sessions, RAG, sandboxing, channels and MCP/A2A/AG-UI, and deploys to AWS, Azure, GCP or Kubernetes via Terraform and Helm.

AI Frameworks | Python · ai-agents · multi-agent

View Project →

yoagent

★ 179

A Rust framework for building tool-using LLM agents, with streaming support across seven protocols, built-in tools, MCP and OpenAPI integrations, sub-agents, and session management. Includes a terminal coding-agent example.

AI Frameworks | Rust · agent loop · tool calling

View Project →