Radar / How scoring works

How JEV scoring works

Every launch on Fresh Weights carries a JEV traction score from 0 (obscure) to 1 (everywhere). It is not a popularity contest and it is not a black box. This page states exactly what is asked, what is fed in, and where the score can mislead.

The short version

How a launch gets its numberOur pipeline reads a launch's primary sources and writes a short evidence file: what it is, why it matters, where it was found, and its GitHub stars. That file goes to JEV, TypeSafe's decision model, with one question. JEV returns a single number between 0 and 1. That number is the traction score you see on every card.

The exact question

Word for word, this is what JEV is asked about every launch:

The question“How popular and noteworthy is this AI launch for a technical, AI-building audience? Consider the launch significance, GitHub stars, and source credibility. Ignore hype and marketing.”

JEV answers typed questions against a state object and returns a structured verdict, not free text. The "noul" question type used here returns a number on a 0 to 1 scale. There is no prompt engineering around it and no second question.

What JEV sees

The state object for a launch has exactly five fields, written by our curation pipeline from primary sources:

title
The launch name, as announced.
summary
One line on what it is, in our words.
usp
Why it matters, in our words.
source
Where it was first found: GitHub, X, Hacker News, a lab account, and so on.
github_stars
Star count at scoring time, or 0 for non-GitHub launches.

What JEV never sees

Vendor marketing copy is not in the state. The summary and USP are rewritten from primary sources, never pasted from an announcement. There is no sponsorship input anywhere in the pipeline, so there is nothing a vendor could buy even if they wanted to. JEV also never sees our opinions: the question is fixed and identical for every launch.

How the score is used

When it is scored

Each launch is scored once, on the day it is added, as part of the four daily update runs. Scores are not re-run later: a launch that blows up a month after listing keeps the score it earned at birth. Treat the number as a first impression from launch week, not a live meter.

Limits, stated plainly

A score is a model's judgment over curated evidence. It is not a measurement of downloads, users, or revenue, and it is comparable across launches without being calibrated to any of those. Three consequences follow:

Read the number like thisThin public evidence scores low even when the launch is good; that is what the “too early to judge” usability verdict is for. GitHub stars can be gamed, and the question compensates by weighing significance and source credibility, but no score built on public signals is immune to gaming. A high score means "worth your attention this week", never "proven in production".

Common questions

Can a launch game its score? It would have to game the underlying evidence: real stars, real builds, real mentions across independent sources. There is no form to fill and no one to email.

Why did a good launch score low? Usually thin evidence at scoring time. The score is a launch-week first impression; the community-builds section on each card is the better trailing indicator.

Which model does the scoring? JEV runs as typesafe/jev-1.13 via OpenRouter. Scoring is budget-capped, so it can never starve the radar's update budget.

Get the week's best AI launches, plus 3 ideas worth building

One email every Saturday. Ranked by traction, not hype. Free.