Review analysis is on every list of things Jev, TypeSafe AI's decision model, is supposed to be good at. We had not seen anyone test it on real reviews, or in any language but English, which TypeSafe itself calls Jev's primary training language. So on 28 September 2026 we collected 840 Google Maps reviews of restaurants, dentists and hotels in three Italian and three American cities and asked Jev two things about them. The first answer was almost boring: it agreed with the stars 99% of the time, in Italian as well as English. The second was more interesting: what the one-star reviews actually blame, and how differently that looks on each side of the Atlantic.
The data
We searched Google Maps for restaurants in Milan and Chicago, dentists in Turin and Austin, and hotels in Naples and Denver, took the first seven businesses of each search, and fetched two pages of reviews for each of the 42: one sorted by lowest rating, one by highest, so that both ends are well represented. That gave 840 reviews. We removed duplicates, reviews under 20 characters, the few three-star reviews, and reviews written in another language than the city's, leaving 730: 346 in Italian and 384 in English, of which 304 had one or two stars.
Collecting them is the routine part. Our Google reviews collector delivers reviews at $0.0005 each, so this whole sample costs about $0.42 to collect again, and the Google Maps collector finds the businesses at $0.001 per place.
Question one: is this review negative?
We sent each review with the business type and one Noul: "overall, the reviewer is dissatisfied with this business". The star rating was never shown to the model; we used it only to score the answers.
| Reviews | Agree with the stars | Positive read as negative | Negative read as positive |
|---|---|---|---|
| English, 384 | 380 (99.0%) | 2 | 2 |
| Italian, 346, instruction in English | 343 (99.1%) | 0 | 3 |
| Italian, 346, instruction in Italian | 343 (99.1%) | 0 | 3 |
On this task there is no language gap: Italian reviews scored as well as English ones, and writing the instruction in Italian changed nothing. That does not contradict TypeSafe's caution about non-English text. Sentiment on a review is about as easy as a judgment gets, and harder questions may well behave differently; test them before relying on them.
The seven disagreements are the most useful output of the whole exercise, because three of them were not Jev's mistake:
- Two Italian hotel reviews with one star and nothing but praise in the text: kind staff, excellent cleaning, a generous breakfast, "we will be back". Jev scored both 0.02.
- One American hotel review with five stars whose text opens by announcing itself as a one-star review and goes on to describe a dirty, poorly maintained room. Jev scored it 0.99.
Those are reviewers who tapped the wrong star. For a business, that is worth knowing: a one-star review full of praise drags the average down and is exactly the one a polite reply can get corrected. One more disagreement was a four-star restaurant review that was mostly disappointment, which is a matter of opinion. The remaining three were genuine misses, all long, mixed reviews that praised one thing and condemned another, with scores between 0.28 and 0.47.
Question two: what do one-star reviews blame?
For the 304 negative reviews we asked a Choice: what is the reviewer mainly unhappy about, among staff, waiting, price and billing, the quality of the core product (the food, the treatment, the room), cleanliness, booking and administration, facilities and noise, or something else.
| Group | Reviews | Top complaint | Second | Third |
|---|---|---|---|---|
| Restaurants, Italy | 60 | Staff 35% | Food quality 30% | Price 20% |
| Restaurants, US | 53 | Food quality 45% | Staff 36% | Booking 8% |
| Dentists, Italy | 27 | Treatment 63% | Staff 30% | Price 4% |
| Dentists, US | 43 | Price and billing 53% | Treatment 30% | Staff 9% |
| Hotels, Italy | 51 | Cleanliness 29% | Room quality 20% | Booking, price, facilities 12% each |
| Hotels, US | 70 | Cleanliness 30% | Facilities 29% | Staff 23% |
The dentists are the striking row. More than half of the American one-star reviews were about money: double charges, surprise bills after insurance, prepaid work. In Italy, where the sample is smaller, almost nobody complained about the bill; they complained about the work itself. Restaurants split the other way on emphasis: Italian reviewers blamed the waiters more often than the food, American reviewers the food more often than the waiters. Hotels were the same everywhere: dirty rooms first.
Two cautions. This is seven businesses per group, from one search each, so read it as what these review pages say, not as a national survey. And these are Jev's labels, so we checked them: of 40 negative reviews drawn at random and read by hand, 30 were labelled as we would have, 5 were reviews with two equally strong complaints where Jev picked the other one, and 5 were wrong: two very short reviews that give no reason at all (the right answer there was "something else"), and three where Jev chose a side complaint over the main one, such as blaming the staff in a review that is mostly about a dirty room. About 75% clear agreement, 88% defensible: good enough for the table above, not good enough to act on a single review without reading it.
What it costs, and how to run it
The negative-or-not question cost $0.0132 for 730 reviews, about $0.018 per 1,000; the complaint question cost $0.0083 for 304, about $0.027 per 1,000. Collecting 1,000 reviews with our collector costs $0.50, so classifying them adds a few percent.
import requests
OR_KEY = "your-openrouter-key"
QUESTIONS = {
"negative": {"type": "noul",
"instructions": "Overall, the reviewer is dissatisfied with this business: the review is negative."},
"complaint": {"type": "choice", "instructions": "What is the reviewer mainly unhappy about?",
"criteria": {"staff": "Rude, unfriendly or unprofessional staff or owner",
"wait": "Waiting, slowness, delays, being ignored",
"price": "Price, billing, extra charges, refunds or money disputes",
"quality": "Quality of the core product or work: the food, the treatment, the room",
"cleanliness": "Dirt, hygiene, smells, pests",
"booking": "Booking, reservations, cancellations, appointments or administrative mistakes",
"facilities": "Noise, location, parking, broken equipment or facilities",
"other": "Something else, or no reason given"}},
}
def analyse(review_text, business_type, stars):
r = requests.post("https://openrouter.ai/api/alpha/decisions",
headers={"Authorization": "Bearer " + OR_KEY},
json={"model": "typesafe/jev-1.13",
"state": {"business_type": business_type, "review_text": review_text},
"questions": QUESTIONS}, timeout=30)
a = r.json()["answers"]
p = a["negative"]["noul"]
wrong_stars = (stars <= 2 and p < 0.1) or (stars >= 4 and p > 0.9)
return {"negative": p, "complaint": a["complaint"]["choice"], "check_stars": wrong_stars}
The check_stars line is the part most review tools do not have: it flags the reviews whose text and stars point in opposite directions with high confidence, which in our sample were all genuine reviewer mistakes.
Where this is useful
- Reputation monitoring. Flag contradictory reviews for a reply, and route complaints by type: billing complaints to the front desk, cleanliness to housekeeping.
- Market research. The complaint mix of a category in a city is a map of what customers care about there. Our table above took minutes and under two cents of model time.
- Lead qualification. A business with a run of recent billing complaints is a different prospect from one with cleanliness complaints, if you sell software or services to them. We tested the qualification side of this in our lead scoring post.
For how Jev compares with LLMs on this kind of judgment, including cost and where their mistakes fall, see our benchmark against five LLMs.