Yesterday I wrote about using Kev to classify Gmail messages locally. The setup worked, but it left a large pile of messages in Not classified.
Today I replaced Kev with Jev and ran the classifier again.
The change is small in code. The effect on the inbox was not.
The new flow
The current pipeline looks like this:
Gmail → gog → OmniRouter → Jev via OpenRouter → gog → Gmail label
gog reads the messages and applies labels. OmniRouter still creates the factual summary before classification. That part stays local on the gateway because long email bodies are a bad input for a small classifier.
The classification step now goes directly to Jev through OpenRouter's System One API. The model is typesafe/jev-1.13.
Jev receives the message metadata and the summary, then chooses one of the existing Gmail categories:
ActionedAwaiting ReplyCold EmailFYIMarketingMeeting UpdateTo Respond
If the result is too uncertain, the message can remain Not classified instead of getting a random label.
Why move away from Kev?
Kev had two practical problems in this task.
First, it was slow on the gateway. The model ran locally on CPU, and some messages took long enough to hit the timeout. Second, its confidence scores were often too low to use comfortably.
I compared both classifiers on a sample of 31 messages that already had historical labels. Kev and Jev chose the same label 19 times, or 61.3% of the time.
Against the historical labels, both reached 22 correct classifications out of 31: 71.0%.
That number doesn't prove that Jev is more accurate. The sample is too small, and the historical labels are not a perfect human ground truth. It does show a clear operational difference:
- Kev's average confidence:
0.322 - Jev's average confidence:
0.869 - Kev had 21 confidence scores below
0.40 - Jev had none below
0.40
Jev also cost about $0.001064 for the 31-message comparison. That's roughly $0.034 per 1,000 messages at the same rate.
The local model was cheaper in the narrow sense that it ran on hardware I already had. But a classifier that spends most of its time timing out, or abstains on a large part of the inbox, isn't necessarily the cheaper system in practice.
What happened to Not classified?
Before the re-evaluation there were 408 inbox messages with the Not classified label.
Those messages were not necessarily broken. They were the messages Kev had left aside because its confidence was below the configured threshold. In the original design that was the safe behavior: when the model isn't sure, don't force a decision.
I ran those messages through Jev in batches. The first batch processed 100 messages:
- 97 received a category;
- 3 stayed
Not classifiedbecause their confidence was still too low; - 0 technical errors occurred.
The remaining batches finished the job. A final Gmail search returned zero messages with the Not classified label.
The labels were distributed across the existing categories, including Marketing, FYI, To Respond, Cold Email, and Meeting Update. I also checked that the Python script still compiled after the changes.
What this result does and doesn't mean
It doesn't mean Jev understands every email correctly. The 71.0% agreement with historical labels is not good enough to call this solved, and those historical labels need a manual review.
It does mean that Jev is more useful for this workflow right now. It gives the script a usable decision more often, costs very little, and avoids keeping a local inference server running just for this one task.
There is also a limit to the Not classified result. Getting the count down to zero is not automatically an improvement if uncertain messages are being forced into the wrong category. The next useful test is a human-labelled reference set, not another impressive-looking count.
For now, Jev handles the first pass and Gmail labels provide the visible result. I will review the disagreements and adjust the prompt, threshold, or category rules where the mistakes are systematic.
The current setup
The classifier still runs automatically through cron. Its credentials come from Bitwarden, including the OpenRouter key, so the key is not stored in the script or the cron entry.
The manual re-evaluation mode is now available as well:
python3 scripts/gmail_kev_classifier.py --reevaluate-not-classified --limit 100
The --limit option makes it possible to process the backlog in manageable batches. That matters because fetching hundreds of complete Gmail messages in one request was unreliable.
The next step is the less exciting one: review the disagreements by hand. That will tell me whether Jev is making better decisions, or simply making more decisions with higher confidence.