Lists (19)
Sort Name ascending (A-Z)
3D生成
文/图生3D、3D mesh生成、PBR材质3D重建与建图
NeRF/3DGS/SLAM/SfM/多视角重建AI Agent
Agent框架、LLM工具调用、MCP、自主代理、ChatUIComfyUI
ComfyUI 本体、节点、插件、工作流世界模型
人脸相关
人脸检测/识别/交换/复原/表情/肖像图像一致性与ID保持
IP-Adapter、角色一致性、ID保持生成图像分割
SAM系列、语义/实例/全景分割图像深度
单目/双目深度估计、深度补全图像生成
Diffusion/GAN 文生图、图生图、风格化生成图像编辑与修复
超分辨率、修复、去噪、抠图、风格迁移、虚拟试衣多模态模型
VLM/MLLM/CLIP/视觉骨干网络大语言模型
LLM训练/推理/微调/RAG/对话工具与基础设施
开发工具、OCR、数据集、教程、VPN、杂项数字人
说话头、数字人对话、人体动作生成、Live2D、VTuber机器人与具身智能
机器人臂、VLA、具身AI、无人机目标检测与追踪
YOLO/DETR/多目标跟踪/点跟踪视频生成与处理
文生视频、视频编辑、视频超分、视频理解语音与音频
TTS/ASR/音乐生成/语音克隆/音频编辑Stars
A Multilingual Translation Model Family
McByte++ : faster, enhanced and long-term tracking version of McByte. With re-ID included.
Detect Anything in Real Time: Real-time object detection using frontier object detection models.
Animate skeletons with natural language; NVIDIA's Kimodo ported to C++/GGML
Official implementation of Kimodo, a kinematic motion diffusion model for high-quality human(oid) motion generation.
[SIGGRAPH Asia 2026] 4DAnyone: Create Anyone in 4D from a Casual Monocular Video
Code for the paper "The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric."
Embodied Passive Aeroacoustic Perception Enables Relative Sensing and Pursuit Between Aerial Robots
SamuraiThirdPersonTemplateThreeJS
AI-assisted PCB design for KiCAD 10. Native KiCAD plugin — a single Rust binary exposing 217 schematic, layout, routing, placement, design-review, and manufacturing tools to Claude, or the LLM of y…
LiteReality-Agent: Turn the real world into simulation-ready environments.
We propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos.
[SIGGRAPH Asia 2026][TOG 2026] InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis
[ECCV 2026 Best Paper Award Candidate] LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction
An open ecosystem of parametric human models and perception stacks, starting with GNM Head.
A new open-source format for interactive video on the web, with a built-in state machine, frame-accurate transitions, and packed-alpha transparency.
An open-weight 11B model series for long-form and real-time video understanding
An Open-Source Project to Unify Audio Processing and Generation
From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models (ReChannel)
[ECCV 2026] SAM2Matting: Generalized Image and Video Matting
Official code for "MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization"
Training open models for agentic phone use with real-app and mock-app environments.
Apple-style Liquid Glass for the web — a headless React lens that refracts the live DOM in Safari, Firefox and Chrome. Zero dependencies.
Zero-dependency canvas shaders that can be installed from npm or designed in Paper



