Open Source
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace. Discover projects across categories like LLM, Vision, Audio, and more.
760 projects
Audio8-AI
A 0.6B Apache-2.0 multilingual TTS model with zero-shot voice cloning, built on a DualAR slow/fast transformer stack, shipping both an INT4 ONNX CPU runtime and an OpenAI-compatible SGLang server.
muscriptor
Kyutai and Mirelo's open multi-instrument music transcription model — audio to MIDI in 103M/307M/1.4B sizes, with a CLI that engraves per-instrument sheet music and tablature via MuseScore.
digimata
A single-binary Swift menu-bar app for macOS that records mic and system audio as separate tracks and transcribes them fully on-device with Parakeet TDT 0.6B v2, getting speaker tagging for free.
cathrynlavery
MIT agent skill giving Claude Code, Codex, and Pi 27 editorial diagram types as self-contained HTML+SVG, with website-driven brand token extraction, WCAG AA contrast checks, and draw.io/Mermaid redraw.
cactus-compute
Cactus Compute's 45M-parameter tool-calling model in a single 14MB binary running in ~28MB of RAM, with grammar-constrained JSON output, built-in tool retrieval, and calibrated confidence gating.
k4yt3x
Long-running C/C++ video super-resolution and frame-interpolation framework wrapping Anime4K v4, Real-ESRGAN, Real-CUGAN, and RIFE over ncnn/Vulkan, with a Qt6 GUI, AppImage, and container images.
antirez
MIT from-scratch C engine running MiniMax-H3 video generation natively on Apple Silicon Metal, with measured speed/quality dials and SSD streaming that cuts DiT memory from 36.5 GiB to 2.0 GiB.
huangruiteng
MIT provider-neutral control plane that keeps objectives, gates, todos, evidence, and quota durable across long-running agent loops on Codex, Claude Code, Cursor, or a custom runner.
VAST-AI-Research
MIT single-image-to-3D-Gaussian model from VAST-AI with an arbitrary Gaussian budget up to 262,144, ~2,000 lines across two files, near-zero dependencies, and official ComfyUI support.
jingyaogong
Apache-2.0 0.1B Omni model trained from scratch that listens, sees, and speaks — Thinker-Talker dual-path with streaming speech and barge-in, runnable end to end on one RTX 3090 in ~2 hours.
AudarAI
Arabic-first generative ASR ranking #1 of 36 on the Open Universal Arabic ASR Leaderboard — a Qwen3 decoder over a Whisper-style encoder, tuned on 300k+ hours across every major Arabic dialect.
studio-dots-ai
Apache-2.0 2B text-to-speech model with no discrete tokens anywhere — continuous autoregressive synthesis over a 48 kHz AudioVAE, reporting SOTA Seed-TTS-Eval scores with a distilled 204 ms-latency variant.