Ask FastVLM about an image
Choose a photo, paste a screenshot or try a sample. The 0.5B model runs on your device and streams its answer here.
Checking browser support…
Your image
Processed on your deviceJPEG, PNG or WebP · up to 10 MB. Paste a screenshot while focused in this panel. Images are resized to a maximum of 1024 px on the longest side.
Choose an image or try a sample
The first run downloads about 1.1 GB from Hugging Face. Later runs may reuse this browser’s cache. Model preparation also needs free GPU memory.
Answer
Choose an image and ask a question.
Timings describe this run on your device. “First text” includes image processing and is not a strict first-token benchmark. JSON includes the question, answer and timings, without image pixels.
Run locally with Python →This image playground processes your image and question locally. Model files are downloaded from Hugging Face and runtime files from jsDelivr. Usage events record action categories and timings, without images, file names, questions or answers. Browser model: onnx-community/FastVLM-0.5B-ONNX.
Try the camera demo
Apple’s separate Hugging Face demo supports a live camera feed. Opening it connects to Hugging Face; camera access is requested inside that demo.
Build this into your own app with the WebGPU starter, or compare the available models and runtimes.
Inspect the original site's 24-image evaluation and failure cases, or benchmark your device.