Open Source
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace. Discover projects across categories like LLM, Vision, Audio, and more.
760 projects
XYZ-AI-Lab
An agentic RL post-training framework wrapping SGLang rollout and Megatron training, focused on multi-turn agent trajectories and on debugging the silent rollout-trainer divergences that destabilize training.
firecrawl
A Rust library that classifies PDFs as text-based or scanned in ~20ms and extracts structured Markdown locally, letting AI document pipelines skip OCR for the majority of files that never needed it.
JustVugg
A pure-C inference engine that runs 744B–2.8T parameter MoE models on consumer hardware by streaming routed experts from disk, treating VRAM, RAM and NVMe as one placement hierarchy.
HUANGCHIHHUNGLeo
Local CLI, MCP server and agent skill that extracts scene-change keyframes with windowed dedup plus a Whisper transcript, so any text-and-image LLM can read a video without a fixed-interval frame dump.
RareSense
Code-native 3D generation that writes Blender Python instead of meshes, producing GLB assets with named parts, an assembly tree and joint pivots that stay editable and riggable after generation.
Kritt-ai
Self-hosted platform that decomposes security review into focused prompts, runs them across Codex or Claude Code agents in parallel, then de-duplicates, validates and ranks the findings.
netease-youdao
NetEase Youdao's Apache-2.0 zero-shot TTS engine clones a voice across 14 languages with no reference transcript, cutting en-to-zh WER to 6.71 against F5-TTS's 11.60.
matthartman
macOS menu-bar dictation and meeting transcription running seven ASR models across WhisperKit, FluidAudio and MLX Audio fully on-device, with a reproducible privacy audit.
wassgha
Transcript-based video and audio editor that runs Whisper, pyannote and ffmpeg entirely in the browser — delete a word, the clip is cut, and the media never leaves the device.
Tencent-Hunyuan
Tencent's Apache-2.0 295B MoE with 21B active parameters and a 256K context, tuned against 50+ product deployments — hallucination rate cut from 12.5% to 5.4% and SWE-Bench variance held within 4%.
facebookresearch
Meta FAIR's Apache-2.0 unified multimodal model that drops both the VAE and the vision encoder for direct pixel patch embeddings — and reports beating the encoder-based variants it was stripped down from.
TencentCloud
MIT-licensed team memory hub from Tencent that turns conversations, docs and code into four governed, reusable agent assets — Chat Memory, Skill, Wiki and CodeGraph — with ACLs and layered L0-L3 distillation.