Everyone showing Jev on lead scoring shows the same thing: a handful of invented leads, a question, a confident number. Nobody shows it on a real list. So we took one. In September we built a list of motor workshops for a lubricant distributor, scraped from a business directory and Google Maps, and cleaning it took several rounds of writing keyword rules. On 28 September 2026 we drew 200 businesses from that list at random, labelled them by hand, and asked Jev, the decision model TypeSafe AI released on 15 September, to sort them with a single sentence. It got 175 of 183 right. Our rules got 176.
A scraped lead list is one third junk
The list started as 7,322 raw rows: every result of searching the directory and Maps for "mechanic workshop" and its synonyms in each town of the province of Caserta, in southern Italy. Merged and deduplicated, that was 1,257 businesses. The buyer sells engine oil and lubricants, so the target is precise: workshops that service or repair cars, vans and trucks. Search does not respect that boundary. Of the 183 businesses we could label with confidence, 63 were not targets:
| What search returned | Businesses |
|---|---|
| Target: repair and maintenance workshops, auto electricians, inspection centres, brand service shops | 120 |
| Tyre shops | 13 |
| Body shops and car glass | 12 |
| Fuel stations (the directory files them under roadside assistance) | 11 |
| Car dealers, brokers and rental | 9 |
| Unrelated: a florist, a bar, a hairdresser, a library, a jeweller, a medical lab, a food company | 8 |
| Parts stores | 4 |
| Industrial machine shops and equipment suppliers | 4 |
| Motorcycle and scooter shops | 2 |
The unrelated ones are there because "officina" means workshop in Italian and turns up in the names of a florist, a hair salon and a cultural centre. Seventeen more businesses in the sample were genuinely ambiguous, for example a listing that Maps calls a tailor and the directory calls a workshop, or a towing company that may or may not repair cars; we left those out rather than guess. Labels were made by one person reading the name, both categories, the directory description and the website domain.
Four ways to sort the same list
Each method saw the same fields a buyer's pipeline would have: business name, Google Maps category, directory category, directory description and website domain. No web visits, no extra data.
| Method | Correct of 183 | Wrong leads kept | Good leads dropped |
|---|---|---|---|
| Keep everything search returned | 120 (65.6%) | 63 | 0 |
| Category filter (the Maps or directory category says workshop) | 173 (94.5%) | 8 | 2 |
| Our keyword rules, tuned by hand on this data for the delivery | 176 (96.2%) | 7 | 0 |
| Jev, one Noul, threshold 0.5 | 175 (95.6%) | 6 | 2 |
The rules deserve a word, because they are the honest baseline. They took several rounds of work: two category allow lists and two deny lists, a name pattern to rescue workshops filed as roadside assistance, and exceptions discovered by reading hundreds of rows (Google labels many workshops in the province "Auto machine shop", which a naive list would drop). They were written on this very data, so they are as good as rules get. Jev got one sentence describing the target, written once, and matched them within one lead.
Where they disagree is instructive. The rules kept a tyre shop because the directory also filed it under workshops; Jev read "professional tyres" in the name and dropped it. Jev dropped a dealer that also runs a service workshop, and a body shop whose name says it is also an auto electrician; the rules, which had been told about those cases, kept both. And both methods kept five body shops that Google files as "Auto repair shop". Those may even be right: many Italian body shops do some mechanical work. They are the cases where a human would pick up the phone.
The probability tells you which leads to check
As in our test on scraped block pages, Jev's probabilities were more useful than its yes or no. Targets scored a median of 0.91; non-targets a median of 0.08. Every one of the eight disagreements with our labels fell between 0.3 and 0.8:
| Policy | Accepted | Dropped | Sent to review | Errors |
|---|---|---|---|---|
| Single threshold at 0.5 | 124 | 59 | 0 | 8 |
| Accept above 0.8, drop below 0.3, review the rest | 109 | 55 | 19 | 0 |
That second line is what a lead vendor wants: 164 of 183 businesses decided automatically with no errors, and 19 (about 10%) flagged for a human, which is roughly one phone call per ten rows. We chose 0.3 and 0.8 by looking at these same scores, so treat them as a starting point, not a law.
What it costs
The 183 calls cost $0.0047 through OpenRouter, a median of 612 input tokens each and 331 ms median latency. That is about $0.026 per 1,000 leads to score. For scale: our Google Maps collector charges $0.001 per delivered place, so filtering with Jev adds under 3% to the cost of the raw list, and the local business leads collector charges $0.01 per delivered lead with contacts. Against a list where a third of the rows are not the buyer's target, a quarter of a cent per hundred leads is the cheapest cleaning step in the pipeline.
import requests
OR_KEY = "your-openrouter-key"
TARGET = ("This business is a workshop that does mechanical maintenance or repair on cars, vans or trucks "
"(an auto repair shop, auto electrician, vehicle inspection centre with a workshop, brand service "
"workshop, radiator, exhaust or gas-system specialist), and so would buy engine oil and lubricants. "
"It is not mainly a body shop, tyre shop, parts store, car dealer, fuel station, towing service, "
"motorcycle or scooter shop, industrial machine shop, or an unrelated business.")
def score_lead(lead):
state = {
"business_name": lead["name"],
"google_maps_category": lead.get("maps_category") or "(none)",
"directory_category": lead.get("directory_category") or "(none)",
"directory_description": lead.get("description") or "(none)",
"website_domain": lead.get("domain") or "(none)",
}
r = requests.post("https://openrouter.ai/api/alpha/decisions",
headers={"Authorization": "Bearer " + OR_KEY},
json={"model": "typesafe/jev-1.13", "state": state,
"questions": {"target": {"type": "noul", "instructions": TARGET}}},
timeout=30)
p = r.json()["answers"]["target"]["noul"]
return "keep" if p > 0.8 else "drop" if p < 0.3 else "review"
How to write the one sentence
The instruction is the whole model, so it is worth writing like a spec:
- Say what the buyer does with the lead. "Would buy engine oil and lubricants" does more work than any list of business types, because it lets the model reason about the edge cases you did not list.
- Name the near misses explicitly. Tyre shops, body shops, dealers and fuel stations are the neighbours search drags in. Naming them is what separates this from "is this a car business".
- Keep structure in code. Phone present, address in the province, business still operating: those are fields, not judgments, and checking them costs nothing.
- Add a Choice when you need the reason. We also asked Jev what each business mainly was; on the 63 non-targets it said tyres 13 times, fuel 11, dealer 9, body shop 7, unrelated 7. That breakdown is what you show a buyer who asks why their list shrank.
What Jev does not replace is the data. It sorted these businesses from five short fields because those fields existed: two independent sources, categories from both, a description where the directory had one. The reason the list was worth sorting is that it was collected from more than one source in the first place. We wrote about what those fields actually contain, and how often, in our analysis of 458 Google Maps leads.