Generate captions and analyze images with various tasks
Classify each person as athlete, coach, official, or other
Compare small VLMs on any image — VQA and field extraction
SAM3 Video Inference on ZeroGPU
sam2 images and video inference on ZeroGPU
SOTA real-time object detection model
pillow bounding box visualization tool
inference for moondream2 point API
Using VLMs for video captioning
Detect objects in images and videos
Identify objects in an image based on text prompts