Jev-X · Building in practice

JEV in Practice: Classifying and Ranking Content in Jev-X

Use one request to classify and score content, then connect the results to filtering, ranking, and display.

Abstract illustration of content streams entering a JEV decision node and connecting to classification, scoring, and ranking steps.
JEV provides judgments; the program organizes the results into a content-curation pipeline.

While building Jev-X, JEV made content classification, scoring, and recommendation ranking easier for me to implement.

Jev-X brings together discussions about JEV on X, helping readers find important updates, tutorials, open-source projects, benchmarks, and practical applications. Collecting posts is only the beginning. The more important questions are: what is each post about, how much useful information does it provide, and which posts deserve to appear first?

Using only keywords and engagement numbers would require me to maintain many rules, and some cases would still be difficult to distinguish. A post containing a GitHub link might announce a tool or simply reference some material. A post with many replies might be useful, or it might have sparked an argument.

With JEV, I can put those judgments directly into the existing pipeline: the model classifies and scores content, and the program uses the results to filter, rank, and display it. For this content-curation project, that lets me build a process that keeps running with relatively little code.

One request, three judgments

For each post that enters the classification pipeline, I ask three questions in the same JEV request, using Choice, Noul, and Score:

Three judgments about one post · scroll sideways to view
JudgmentTypeQuestionResult used by the program
ClassificationChoiceWhat is the post mainly about?Category and category confidence
RelevanceNoulIs JEV or TypeSafe AI the main subject?Relevance probability
Content qualityScoreHow informative is it for someone following JEV?Quality score
Diagram with Chinese labels: the post and its context enter one request; Choice returns a category and confidence, Noul returns relevance probability, and Score returns a quality score from 0 to 4.
Chinese diagram: one request handles classification, relevance, and quality scoring. View full-size image ↗

The input is the content and its context. The output is a structured result the program can use directly. TypeSafe’s official documentation supports combining these questions in a single request. [1]

The category can become a filter label, while the quality score helps determine whether a post enters the curated selection and how it contributes to ranking. How well the prompts fit the project determines how useful those results are.

Classification: define the boundaries clearly

The classification prompt defines nine categories: integration, open source, comparison, tutorial, showcase, benchmark, interview, news, and opinion. Each category has a definition, exclusions, and examples.

For example, a platform announcing support for JEV usually belongs under “integration.” Someone showing an application they built with JEV usually belongs under “showcase.” A repository, CLI, or plugin released for others to install is a better fit for “open source.” All might mention building with JEV, but readers look for them for different reasons.

I ask the model to classify a post by its main purpose. If several categories fit, it should choose the one a reader would most want to use as a filter.

“Tutorial” and “benchmark” also have explicit criteria. A tutorial must teach something or clearly introduce a guide the author wrote. A benchmark must contain results from actual testing. Repeating an official speed claim does not automatically make a post a benchmark.

When category confidence is too low, the program falls back to keyword rules. JEV reduces the amount of judgment I need to encode manually, while the existing pipeline keeps this fallback.

Relevance: record and observe first

The relevance question separately checks whether a post seriously discusses JEV, merely mentions it or adds a hashtag, or refers to a person or another product with the same name.

The current default configuration saves this result for observation and adjustment. It does not use relevance as a mandatory filtering threshold.

Quality: content first, engagement second

The most important instruction in my quality prompt is: judge what the content actually delivers, rather than how excited it sounds.

Instead of simply asking the model for a score from 0 to 100, I define five understandable levels:

  1. Noise: spam, giveaways, engagement bait, or unrelated content.
  2. Reaction only: emotion, hype, or a one-line take with no information.
  3. Some information: an opinion with reasoning, or a short summary of known facts.
  4. Concrete information: real news, integration progress, a demo, measured results, or a clear explanation.
  5. Worth returning to: original insight, a reusable method, a tutorial, an open-source tool, or data worth saving, with reader engagement as supporting evidence.

JEV returns a numeric score from 0 to 4 using these criteria. The program then converts it to the 0–100 score displayed on the page. It represents the content’s value to the intended reader.

What goes into the input

To provide enough context, I include the text, author, quoted content, links, media type, and engagement data. Engagement includes likes, reposts, quotes, replies, bookmarks, and views, along with calculated engagement metrics.

The prompt asks the model to judge the content first and use engagement as supporting evidence. A high bookmark-to-view ratio might suggest readers want to save a post. Giveaways, “reply 1,” follow-and-repost requests, and arguments should not earn a high score just because they attract engagement. Popularity alone is also insufficient for the highest quality level.

For posts introducing articles, repositories, or videos, I ask the model to judge their value from the description in the post. A short post that clearly explains what a tool does can still be useful.

The current pipeline does not have JEV read full linked pages or watch video footage. It judges the text and metadata supplied to it. That is an important boundary when interpreting a score.

Ranking: put quality into the formula

Once the quality score is available, the program still handles the final ranking. The current formulas are simple:

Engagement score = likes × 0.5 + (reposts + quotes) + replies × 13.5

Bookmarks and views enter the prompt, but do not directly enter this engagement formula.

Ranking score = JEV quality score × log10(10 + engagement score)

Taking the logarithm of engagement compresses the numerical advantage of popular posts, giving quality a meaningful role in the ranking.

Consider two hypothetical posts:

  • Quality 80, engagement 100:
    80 × log10(110) ≈ 163.3
  • Quality 25, engagement 1,000:
    25 × log10(1010) ≈ 75.1
Chinese diagram of a hypothetical comparison: quality 80 and engagement 100 produce a ranking score of about 163; quality 25 and engagement 1000 produce about 75.
Chinese diagram: the higher-quality post ranks first in this hypothetical example. Both quality and engagement affect the actual ranking. View full-size image ↗

This example shows that a higher-quality post can outrank one with more engagement. The formula still considers both, so it does not guarantee that every higher-quality post will outrank every more popular post.

The curated selection also applies engagement thresholds, candidate limits, low-quality filtering, and limits on posts from the same author. This is content curation without user profiles or personalized recommendations.

The pipeline: from collection to display

Chinese diagram of the six-step Jev-X pipeline: collect X posts, save to D1, call JEV through a queue, save categories and scores, filter and rank into KV, then read and display on the website.
Chinese diagram: collection, model judgment, and display each have a clear role in the pipeline. View full-size image ↗
  1. Collect relevant X posts
  2. Save the content to D1
  3. Call JEV through a queue
  4. Save categories and scores
  5. Filter, rank, and write to KV
  6. Read the results on the website

Newly collected content that enters the classification scope is registered as a task. Once JEV’s results have been saved, the backend organizes and publishes the lists.

If a post already has complete results for the current prompt version, those results can be reused. When engagement numbers change, the program recalculates the ranking during a later publication without calling the model every time. Saved judgments can also be reused after a publication failure.

The website reads the prepared results directly, so readers do not wait for the model to score posts when they open a page. JEV provides one explicit, reusable judgment step.

Cost and speed

Low cost is another reason I am happy to keep using JEV. As of October 2, 2026, TypeSafe’s official model documentation lists a price of $0.042 per million input tokens, with output tokens free. [2]

For an illustrative estimate, suppose one post and its prompt consume 2,000 input tokens. The model cost would be about $0.000084 per request, or $0.84 for 10,000 posts.

This calculation uses an assumed input length; it is not a measurement of Jev-X’s token usage. Actual cost varies with the text and prompt length, and excludes X data collection and infrastructure. That price level makes classification and scoring practical as part of my everyday pipeline.

Speed matters too. TypeSafe’s launch announcement reports an end-to-end response time of 70–500 milliseconds. [3] This is the official figure. Jev-X has not recorded its own latency distribution, and the figure is not the total time from collecting a post to displaying it on the website.

For a project that repeatedly classifies and scores content, low cost and short response times directly affect how practical it is to keep the process running.

Making content judgment an everyday step

Using JEV in Jev-X has made its value more concrete to me: it makes content judgment easier to build into software.

I still need to define categories, design scoring criteria, handle low-confidence results, and check what actually appears on the page. Scores are a guide for reading and filtering; they cannot replace understanding the original content.

But with JEV, I can spend more of my attention on product decisions: what readers need, which posts should appear first, and how to make useful experiments easier to discover.

Make classification and scoring an affordable everyday step, then use simple, clear code to turn the results into content readers can browse. That is the practical help JEV has brought to this project.