Blog / Questions / FIG. 51
Can Jev Judge Images? Not Directly, and the Workaround Is the Lesson
Jev judges text, not pixels. But the screenshot-free pattern, local extraction feeding text verdicts, handles more image work than you'd expect.
Directly, no: Jev is a text model, per everything the ecosystem documents. Feed it pixels and you have a category error; there's no vision pathway to prompt around. But the practical question underneath, "can a cheap verdict layer participate in image-heavy workflows?", has a much better answer, and one cataloged build turned it into the pattern this page exists to teach.
The screenshot-free computer-use build drives a computer with no vision model at all: local OCR turns the screen into text, and Jev rules on the text, every decision, at verdict prices, with raw pixels never leaving the device. Generalize it and you get the template: cheap local extraction, remote text verdicts. Documents become judgeable after OCR; product images become judgeable via their listings, alt text, and extracted labels; screenshots in support tickets become judgeable through what's readable in them; even moderation pipelines can route text-in-image violations (the classic filter dodge) by OCR-ing first, per the spam page's arms-race logic. The division of labor is clean: perception stays local and specialized, judgment goes to the fast generalist, and your privacy posture improves as a side effect because pixels never travel.
Where the workaround honestly ends: judgments about the image as image, aesthetics, layout quality, whether the photo shows a counterfeit, what's happening in a scene. Those need a vision model, full stop, and the right architecture is the usual cascade with a vision tier: extract-and-judge for the text-representable 80 percent, vision model for the genuinely visual slice, humans on consequence. The ecommerce counterfeit case from that vertical's page is the canonical example of knowing which side of the line you're on.
Frequently asked questions
Will Jev add image support?
No public roadmap says so; the text-only shape looks like design, not gap. Build the extraction pattern and you're insulated either way.
What OCR should I pair with it?
Whatever runs where your data lives: on-device or in-VPC extraction preserves the privacy win that makes this pattern shine. The build that started it used local OCR deliberately.
Can it judge video?
Same answer one level up: not the frames, but transcripts and extracted text absolutely, which is how the call-QA pipelines already work.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.