UniSpace
Unified text-to-image generation and instruction editing
None defined yet.
Unified text-to-image generation and instruction editing
Control-video + prompt to video with sound, MiniMax-H3
Causal world model for robot video generation
Omni-modal image, audio and video understanding
4-step MiniMax-H3 β video with a matching soundtrack
Anime text-to-image with a Qwen3.5 4B cross-adapter
Japanese streaming ASR fine-tuned on 35k hours of speech
TinyCast zero-shot probabilistic time-series forecasting
Generate human motion sequences from text prompts
Monocular portrait video to multi-view videos
Unified text-to-image and image editing model
Action-conditioned robot manipulation video generation
Track fish in sonar video, fix it with plain English
Multilingual word and sentence alignment
Detect AI-generated images with FiSeR DINOv3 ViT-L/16
Blind image quality assessment with faithful reasoning
Assess spatial aesthetics of interior images
Unified text-to-image and image editing (SFT checkpoint)
Krea 2 Turbo in 4 steps via distillation LoRA
4-step distilled image-to-video with SLA sparse attention
Streaming video QA from the 4 most recent frames
Chinese dialect ASR for Mandarin, Cantonese, Wu & more
Instruction-based video editing with Qwen-Image-Edit
Turn natural-language math into Lean 4 statements