Open Source
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace. Discover projects across categories like LLM, Vision, Audio, and more.
760 projects
bytedance
ByteDance's Apache-2.0 unified video generation/editing framework pairing a 7B MLLM semantic planner with a 14B DiT renderer — first-tier against commercial editors in blind human-preference arena voting.
debpalash
The open-source ElevenLabs alternative: local zero-shot voice cloning, video dubbing, dictation, and audiobook production across 646 languages, orchestrating 14 TTS + 10 ASR engines on your own hardware (AGPL-3.0).
kyutai-labs
Kyutai's MIT-licensed 100M-parameter TTS that runs ~6x real-time on 2 CPU cores with ~200ms latency — streaming audio, voice cloning from a WAV, 6 languages, and zero GPU required.
unclecode
The most-starred open-source web crawler on GitHub (Apache-2.0): turns any site into clean, LLM-ready Markdown for RAG and agents, with async crawling, structured extraction, and a secure-by-default Docker server.
mozilla-ai
Mozilla.ai's Apache-2.0 project that collapses an LLM and its runtime into a single cross-platform executable — built on llama.cpp and Cosmopolitan Libc, running install-free on macOS, Linux, Windows, and BSD.
pytorch
PyTorch's official BSD-3 platform for large-scale LLM pretraining — a clean, minimal reference codebase composing FSDP2, tensor/pipeline/context parallelism, torch.compile, and Float8/MXFP8 low-precision training.
HKUDS
HKUDS's open-source, MCP-native personal trading agent (MIT) pairing a provider-agnostic multi-agent runtime with a 460+ factor Alpha Zoo, hardened backtesting, and multi-market broker execution.
TencentARC
Tencent ARC Lab's SIGGRAPH 2026 single-image-to-3D model (MIT) that back-projects pixel features into 3D for high-fidelity GLB meshes with PBR textures, shipping inference, demo, and training code.
zai-org
Z.ai's open-weight flagship vision-language model (Apache-2.0), a hybrid-reasoning MoE VLM with grounding, GUI-agent, and video-understanding capabilities, shipped in standard, FP8, Flash, and GGUF variants.
thewh1teagle
An open-source, privacy-first desktop app (Rust, MIT) that transcribes audio and video fully offline using Whisper, Parakeet, and Nemotron, with diarization, many export formats, and built-in AI summaries.
KittenML
KittenML's ultra-lightweight open-source TTS: ONNX models from 15M to 80M params (as small as 25 MB) that synthesize high-quality speech on CPU with 8 built-in voices (Apache-2.0).
FunAudioLLM
Alibaba FunAudioLLM's open-source Any2Audio framework (NeurIPS 2025) that generates and edits video/text/audio-conditioned sound using Chain-of-Thought reasoning from MLLMs.