Open Source
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace. Discover projects across categories like LLM, Vision, Audio, and more.
760 projects
QwenLM
Apache-2.0 vision-language model family from Alibaba's Qwen team spanning 2B dense to 235B MoE, with native 256K context, GUI agent capability, and 3D grounding — the cleanest open VLM choice for edge-to-cloud deployments in 2026.
OpenWhispr
MIT-licensed cross-platform voice dictation and meeting transcription app with local Whisper/Parakeet models, BYOK cloud LLMs (GPT-5/Claude/Gemini/Groq), on-device speaker ID, and an MCP-integrated voice agent.
KoljaB
MIT-licensed low-latency Python library for real-time speech-to-text with WebRTC/Silero VAD, Porcupine/OpenWakeWord wake words, FastAPI streaming server, and pluggable engines (faster_whisper, Moonshine, sherpa-onnx).
FunAudioLLM
MIT-licensed multilingual speech understanding foundation model: ASR + emotion + audio event detection + language ID across 50+ languages, 15x faster than Whisper-Large with non-autoregressive inference.
lfnovo
MIT-licensed self-hosted NotebookLM alternative supporting 18+ AI providers, multi-modal source ingestion (PDFs, YouTube, audio, Office), multi-speaker podcast generation, hybrid search, and a full REST API.
NVIDIA
OpenMDW-1.1 licensed open platform from NVIDIA pairing a Reasoner (text and vision in, text out) with a Generator (multimodal in, video, audio, and action sequences out) for robotics, autonomous vehicles, and physical AI.
NousResearch
MIT-licensed self-improving AI agent from Nous Research with persistent skill learning, 200+ model compatibility, 6 execution backends, and integrations for Telegram, Discord, Slack, WhatsApp, and Signal.
chopratejas
Apache-2.0 Python and TypeScript project that compresses tool outputs, logs, files, and RAG chunks before they reach the LLM. Ships as a library, drop-in HTTP proxy, and MCP server with documented 60-95% token reductions on real agent workloads.
lyogavin
Apache-2.0 Python library that runs 70B and even 405B LLM inference on a single 4GB GPU using layer-wise sequential loading from disk, with optional 4-bit and 8-bit block-wise quantization for speed.
Open-LLM-VTuber
Local, multimodal voice agent that combines swappable ASR, LLM, and TTS backends with a Live2D avatar and voice interruption, running fully offline on Windows, macOS, and Linux.
roboflow
Reusable computer vision tools for detection, tracking, and segmentation pipelines.
supertone-inc
Lightning-fast on-device multilingual TTS engine running natively via ONNX.