Open Source
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace. Discover projects across categories like LLM, Vision, Audio, and more.
760 projects
alphacep
Offline open-source speech recognition toolkit with 20+ languages, tiny 50MB models, a streaming API, and bindings for Python, Java, Node.js, C++, Rust, and Go.
m87-labs
Tiny open-source vision language model (2B and 0.5B) for captioning, visual Q&A, and object detection that runs efficiently on-device or in the cloud under Apache-2.0.
resemble-ai
State-of-the-art open-source TTS from Resemble AI with multilingual voice cloning, a low-latency Turbo model, and native paralinguistic tags — all under an MIT license.
altic-dev
Open-source macOS voice-to-text dictation app that transcribes fully on-device, supporting Nemotron Speech, Parakeet, Whisper, and Apple Speech with local AI enhancement.
black-forest-labs
Black Forest Labs' open-weight image generation and editing models, with a minimal official inference repo spanning text-to-image, inpainting, structural conditioning, and Kontext editing.
THUDM
Open-source LLM post-training framework for RL scaling that connects Megatron training with SGLang rollout in a single, battle-tested loop behind the GLM models.
topoteretes
Open-source AI memory platform that gives agents persistent long-term memory by ingesting any data into a self-hosted knowledge graph combining vector embeddings and graph reasoning.
kvcache-ai
A flexible framework for cutting-edge LLM inference and fine-tuning via CPU-GPU heterogeneous computing, letting huge MoE models like DeepSeek and Kimi-K2 run on limited GPU memory.
alibaba
Alibaba's lightweight in-page GUI agent that controls web interfaces with natural language using text-based DOM manipulation — no browser extension, headless browser, or multimodal LLM required.
TencentARC
Tencent ARC's open feed-forward framework that turns a single image into a textured 3D mesh in about ten seconds using sparse-view large reconstruction models.
QwenLM
Alibaba's flagship open vision-language model series with visual-agent UI control, image-to-code generation, 3D grounding, and native 256K context expandable to 1M tokens.
m-bain
A fast open ASR system that wraps Whisper to add accurate word-level timestamps, 70x real-time batched inference, and speaker diarization for subtitles and meetings.