🔍 Read the full analysis: 24 Jev Ideas For Modeling Decisions In AI on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer’s Sept. 29 article maps 24 potential uses for Jev, a tool that returns typed answers to narrow questions so software can make small decisions. Meyer says three uses are live, 12 are strong fits and seven need measurement; two are poor fits. The account is based on his own operation, and the article does not provide independent validation of its performance claims.
Thorsten Meyer published a 24-use map for Jev on Sept. 29, reporting that three applications are already running in his publishing operation, 12 meet his criteria for a strong fit, seven need measurement and two are poor fits. The assessment sets out criteria for evaluating where automated judgments may be used and identifies applications Meyer says still require measurement.
Meyer describes Jev as a system that receives text or JSON plus typed questions and returns answers that software can act on. Its answer types include a yes-or-no probability, a choice among options, or a score on ordered levels. The system does not write or summarize the material, according to the article; the surrounding code determines what to do with each response.
The three live examples are a relevance check for stories against a site profile, a language check, and a fallback topic classifier. Meyer says a scan of 78,889 articles cost $2.01, found 1,576 non-English articles and fixed 1,553. He also reports roughly 10,000 relevance pairings judged in three days, with 22% clearly on-topic, and 89% agreement with a frontier language model for the classifier overall.
For that classifier, Meyer reports 97% to 99% agreement when Jev’s confidence was at least 0.8, compared with 42% below 0.5, in a 31-topic measurement. These are results reported by the author about his own operation; the article does not detail an independent replication or all evaluation methods.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Small Decisions May Pay Off
Meyer proposes using Jev for large volumes of narrow judgments where software can act on typed answers. Under his approach, code applies a rule when the answer is clear, while uncertain cases can be referred to a person or a more capable model. He says this design could support checks across many items while routing uncertain answers for further review.
His examples vary in readiness: he rates disclosure checks and comment moderation as strong publishing fits, while a thin-source detector and headline-quality score need measurement first. He classifies a duplicate-story detector as a poor fit because his canary found no duplicates. The article’s criteria include assessing whether a proposed use addresses a measured need.
The Four Conditions Behind the Ratings
Meyer says a suitable Jev task has high volume, a narrow question, cheap errors or a path for uncertain cases to a stronger system, and a heuristic that demonstrably fails. He advises keeping a keyword rule when it works and testing a replacement in shadow mode before switching it on.
His proposed test is to replay 300 to 500 past decisions, compare results overall and by confidence band, and inspect 20 disagreements. He says integration should proceed only where the high-confidence band reaches 95%. The article also recommends an off-by-default feature flag and a canary on 5% to 10% of units before wider rollout.
The article begins its catalog with publishing and content, giving six examples and describing three in detail in the supplied material: source sufficiency, product relevance in roundups, and headline quality among the candidates needing measurement; disclosure checks and comment moderation are rated strong fits, while same-event deduplication is a poor fit. The source text provided for this report ends as the commerce and customer-operations section begins, so it does not establish the full list of remaining applications.
“Jev does not write, summarise or extract. You send it a state (text or JSON) and a set of typed questions, and it returns calibrated answers your code can branch on, with no prose to parse.”
— Thorsten Meyer, describing Jev
Evidence Still Depends on Meyer’s Tests
The reported costs, volumes and agreement rates come from Meyer’s account of his own publishing operation. The material does not identify an independent evaluator, provide full test data, or explain enough about the comparison model and sample selection to assess how broadly the results apply. The 97% to 99% figure applies specifically to cases with confidence of at least 0.8 in a 31-topic classification measurement; it should not be read as overall accuracy across all Jev tasks.
The supplied article text is incomplete after the start of the commerce and customer-operations section. It says the catalog covers publishing, commerce, software, business operations and the home, but it does not show the rest of the 24-use list. The number and details of the applications in those later categories therefore cannot be confirmed from the available material.
Measure Before Wider Deployment
Meyer’s recommended next step for prospective users is to test a candidate task against real past decisions, inspect disagreements, and check performance within confidence bands. A team that meets his proposed threshold can then try the change behind a feature flag, beginning with a small canary and sending uncertain cases through the existing process.
The article does not announce a product launch, independent benchmark or rollout date. Whether the listed candidates outside Meyer’s live publishing workflows meet the test remains to be measured by teams applying them to their own data.
Key Questions
What is Jev, according to the article?
Jev returns typed answers to narrow questions about text or JSON. Software can use those answers to route, classify, filter or score items.
How many of the 24 proposed uses are already running?
Meyer says three are live in his publishing operation, 12 meet his strong-fit criteria, seven need measurement and two are poor fits.
What does Meyer recommend before using Jev in production?
He recommends replaying 300 to 500 real past decisions, comparing performance across confidence bands and reviewing disagreements. He proposes wiring in a use only where the high-confidence band reaches 95%, then starting with a small canary.
Are the reported performance figures independently verified?
The supplied material attributes the figures to Meyer’s measurements. It does not describe an independent replication.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
