Open Source
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace. Discover projects across categories like LLM, Vision, Audio, and more.
760 projects
Panniantong
Open-source toolkit that gives any AI agent free, configuration-light access to read and search the internet — Twitter, Reddit, YouTube, GitHub, Bilibili, and more.
Qualcomm
High-performance on-device inference framework running the latest multimodal models on NPU, GPU, and CPU across Android, Windows, and Linux — often with day-0 model support.
calesthio
The first open-source, agentic video production system that turns your AI coding assistant into a full studio — research, scripting, asset generation, editing, and composition.
microsoft
Microsoft's large 3D asset generation model (CVPR'25 Spotlight) that turns text or image prompts into high-quality 3D via a unified Structured LATent (SLAT) representation, decoding to Radiance Fields, 3D Gaussians, or meshes. Up to 2B params, MIT-licensed.
m87-labs
A tiny open-source vision language model from m87-labs that runs anywhere, offering captioning, visual question answering, and object detection. Ships a 2B general model and a 0.5B edge-optimized variant under Apache-2.0.
SYSTRAN
A CTranslate2-based reimplementation of OpenAI's Whisper that runs up to 4x faster at the same accuracy with lower memory, adding 8-bit quantization, batched inference, and word-level timestamps. MIT-licensed and FFmpeg-free.
resemble-ai
Resemble AI's MIT-licensed family of state-of-the-art open TTS models, spanning a 350M low-latency Turbo variant and a 23+ language Multilingual V3. Offers zero-shot voice cloning, exaggeration control, and built-in Perth neural watermarking on every clip.
multimodal-art-projection
An open-source, LLM-based foundation model for full-song music generation — a transparent alternative to Suno. A 7B stage-1 model plus a 1B refinement stage turn lyrics into complete songs with coordinated vocals and instrumentals across five languages, under Apache-2.0.
facebookresearch
Meta AI's foundation model for promptable segmentation in images and video, using a streaming memory module to track objects across frames from a single click, box, or mask prompt. Ships Tiny-to-Large checkpoints under Apache-2.0.
jd-opensource
JoyAI-Image is JD.com's open-source unified multimodal foundation model that combines an 8B Multimodal LLM with a 16B Multimodal Diffusion Transformer to handle image understanding, text-to-image generation, and instruction-guided spatial editing inside a single closed-loop architecture. It introduces the OpenSpatial-3M dataset and SpatialEdit benchmark for controllable image manipulation.
zai-org
GLM-V is Z.ai's open-source vision-language model series — spanning GLM-4.1V-Thinking, GLM-4.5V, and GLM-4.6V — that unifies image and video understanding with native multimodal function calling, a 128K token context window, and scalable reinforcement learning for complex visual reasoning. The 106B flagship and a 9B Flash variant are both fully open-source.
QwenLM
Qwen3-Omni is Alibaba Cloud's natively end-to-end omni-modal LLM that jointly processes text, images, audio, and video and generates real-time streaming speech and text responses. It achieves state-of-the-art results on 22 of 36 audio/video benchmarks and open-source SOTA on 32 of 36.