Blog / 26
Jev for Marketers: When Growth Decisions Stop Being Vibes
What the content-and-growth builds prove: ad teardowns at scale, pre-publish scoring, reply filtering, and marketing judgment as cheap measurable verdicts.
Marketing runs on judgments that everyone pretends are analysis: this ad feels strong, that subject line seems off, this reply looks like a real prospect. The judgments aren't wrong, exactly; they're unpriced, unlogged, and unrepeatable, which means they can't compound. The content and growth category is where builders started converting those judgments into structured verdicts with Jev, TypeSafe AI's decision model, and it's quietly one of the most commercially instructive corners of the directory. Numbers below are as reported by build authors, per house rules.
Competitive teardowns at absurd scale
The flagship receipt: one builder ran 724 competitor ads through a judgment battery in about 40 seconds for $0.09 (build). Per ad: does it lead with price, fear, or aspiration; does it name a competitor; does it make a verifiable claim; who's the implied audience. That's a competitive-intelligence report that agencies bill weeks for, produced between meetings, and re-runnable monthly so it becomes a trendline instead of a binder. The general recipe is the judge pattern pointed at anyone's public output: competitors' ads, landing pages, email sequences, app-store reviews.
Pre-publish scoring: the editor who's always in
The SuperX scorer asks 61 questions per draft in about a second for $0.0004: is there a concrete claim, does the hook survive the first line, is there a reason to reply. Whatever your channel, the shape ports: a house style guide rewritten as operational judge questions becomes an instant pre-publish gate, and unlike a human editor, it scores every draft, including the 6 a.m. ones. The craft warning applies double in marketing: "is this engaging?" is astrology; "does the first sentence contain a specific number or named entity?" is a ruling.
Inbound and reply triage: the revenue edge
Growth generates streams, and streams are triage: campaign replies (interested, objection, unsubscribe, out-of-office, each to its own workflow), comment sections filtered for substance versus reply-guy noise (a cataloged build does exactly this), and lead routing, where the reported reference is 700 leads scored in 40 seconds for $0.09 (the full lead-scoring guide). This is the least glamorous section and the one your CFO will like: reply-speed on real prospects is the most reliably purchased improvement in outbound, and it's bought here for cents.
Simulation and research, with a seatbelt
The most avant-garde build in the category rebuilt a feed-ranking algorithm to simulate how content might perform pre-publish, virality as a judged property rather than a lottery ticket. We catalog it with the same eyebrow we give the trading experiments: directionally fascinating, and a model's verdict about engagement is a hypothesis about humans, not a measurement of them. Use simulated scores to rank drafts against each other, never to promise outcomes; then let the actual audience file the binding ruling. Same honesty applies to research-style batch judging of survey text, reviews, and social listening: judged piles beat unread piles, and neither beats talking to customers.
Frequently asked questions
How can marketers use a decision model without engineering help?
Start with a batch job on an export: dump competitor ads or campaign replies into a sheet, run a question set, read the verdict columns. The getting-started guide is written for exactly this afternoon; live integrations can wait until the offline pass earns them.
What does marketing-scale judging cost?
Reported references: 724 ads for $0.09, 700 leads for $0.09, per-draft scoring at $0.0004. Entire-channel audits price below one stock photo; the cost table has the wider context.
Can it write my ads or posts?
No; Jev doesn't generate. It judges drafts, teardowns, and streams. Pair it with your writer, human or frontier model, and let the verdicts pick which draft ships.
Is AI scoring of content actually predictive?
It's consistent, which vibes aren't, and consistency lets you learn. Treat scores as a ranking signal validated against your own performance data over time, per the evals mindset, not as a crystal ball.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.