BTInsightsBTInsights

Jev vs. OpenAI Decisions API: Which Codes Open-Ends More Accurately?

OpenAI’s Decisions API coded open-ended survey responses slightly more accurately than TypeSafe’s Jev, but only after we adjusted how strict each model was. With default strictness settings, the two were equally accurate, however Decisions cost about 7 times more than Jev.

At BTInsights, we tested both models on 2,617 real survey responses that human coders had already coded. Then we measured how often each model’s codes matched the human coders’ codes.

Neither model is designed for coding survey open-ends. Both are decision tools built to make fast, low-cost structured calls on any kind of text. Platforms built for survey coding, such as BTInsights, code open-ends much more accurately.

Abstract cover art: soft blue and green textured light with the words "Jev vs. OpenAI Decisions" in white.

Key takeaways

  • Equally accurate with default settings. On an accuracy score out of 100, Jev scored 72.2 and OpenAI Decisions 72.9, a difference too small to matter.
  • Decisions pulls slightly ahead once adjusted. After we adjusted how strict each model was, Decisions scored 77.1 and Jev 75.4.
  • The winner depends on the question. Decisions was more accurate on a question about pets, Jev on a question about urgent care. They tied on a question about AI.
  • Both apply too many codes with default settings. About a third of the codes each model applied didn’t match the human coder’s. Making the models stricter fixed much of that.
  • Decisions costs about 7 times more than Jev.

Coding comparison terms in plain English

  • Accuracy score. How closely a model’s codes match a human coder’s, on a scale of 0 to 100. A score of 100 means a perfect match. Researchers call this the F1 score; we explain how it works below.
  • Confidence score. These models don’t answer a plain yes or no. For each code, they say how sure they are that it fits, such as “87% sure”.
  • Strictness setting. How sure a model must be before it applies a code. The default for both models is 50%: apply the code if the model thinks it’s more likely to fit than not. “Adjusted” means we found a better setting for each question, using responses a human had already coded.

Jev vs. OpenAI Decisions at a glance

JevOpenAI Decisions API
Made byTypeSafeOpenAI
Accuracy with default strictness (out of 100)72.272.9
Accuracy after adjusting strictness (out of 100)75.477.1
More accurate onUrgent care questionPet question
Relative cost in our test1×About 7×
Question formatsYes/no, pick one, ratingYes/no, pick one, rating
Status (October 2026)AvailablePublic beta

What are Jev and the OpenAI Decisions API?

Jev is a decision model from TypeSafe. The Decisions API is OpenAI’s equivalent. You give either one a piece of text and a question, and they return a structured answer with a confidence score, such as “92% sure the answer is yes”. They don’t write, summarize or follow long instructions. In exchange, they are fast and cheap.

This makes them ideal for high-volume jobs where speed and cost matter more than nuance. OpenAI gives examples of expected use: classifying content, routing requests and prioritizing work.

Both models offer the same three question formats:

FormatThe question it answersJevOpenAI Decisions
Yes/no“Does this text fit X?”NoulPredicate
Pick one“Which of these options fits best?”ChoiceChoice
Rating“Where does this fall on a scale?”ScoreScore

For this test, we compared Jev and Decisions using the yes/no format, asking about one code at a time and also compared these to a human coder.

The dataset: 2,617 human-coded survey responses

A fair accuracy test needs a reference answer that you can trust. Coding open-ended survey responses, also called verbatim coding, gives us one. Human coders had already read every response and tagged them with codes from a codebook, so we could check each model’s codes against theirs.

We used three open-ended questions from a consumer survey run by ANR, a market research firm, with 2,617 responses in all.

QuestionResponsesTypical answer
“Tell us what you love most about your pet(s).”602A phrase or a few sentences
“In what ways are hospital-affiliated urgent care centers different from urgent care centers that are not affiliated with a hospital or health system?”983A phrase or a few sentences
“What one word best describes how you feel about Artificial Intelligence (AI)?”1,032One word

Here are a few responses and the codes a human coder gave them:

QuestionResponseHuman coder’s codes
Pets“Always happy to see you!”Unconditional love/happy to see me
Pets“Comfort. Provide a purpose. Makes the house feel less empty now that kids have moved out.”Calm/caring/support/comfort; Responsibility/need/purpose
Urgent care“All your health information is linked.”Access of records/continuity of care
AI“Afraid”Negative – concern/fear/distrust
AI“Ambivalent”Neutral – mixed emotions

Each question was coded twice, by two human coders using their own codebooks of 8 to 14 codes. That gave us six reference answers. Catch-all codes such as “Other” and “Don’t know” were not scored.

How the test worked

The models give a confidence score, while a human coder gives a plain yes or no. Here is how we turned one into the other and compared them.

Flowchart of how BTInsights tested the coding accuracy of Jev and the OpenAI Decisions API against human coders

Step 1: The model scores every code

For each response, we asked each model about every code in the codebook, for example: “Does this response fit the code ‘Loyalty’?” It answered with a confidence score. 95% means “almost certainly yes”, 3% means “almost certainly no”, and 50% is a coin flip.

Step 2: The strictness setting turns scores into yes or no

A code is applied only if the model’s confidence reaches the configured strictness setting. By default, the strictness setting is 50%. Raise it, and the model must be more confident, so it adds fewer wrong codes. But raise it too far, and it starts missing codes that do fit.

Step 3: Adjusting the strictness setting

50% isn’t always the right bar. To find a better one, we took a few hundred responses a human had already coded, kept separate from the test responses. We tried a range of settings on them and kept the one that best matched the human.

It’s like onboarding a new human coder. They review work your lead coder already did, learn whether to be stricter, then start the real job.

Step 4: Compare with the human

Both models coded the same 250 test responses per question, first with default strictness settings, then again with our adjusted settings. We compared their codes with the human coder’s, then repeated the check on all 2,617 responses.

An example

One pet owner wrote: “Always happy, great cuddles, make us laugh, unconditional love, teaches kids about responsibility and keep us active.” The human coder applied four codes. Here is how sure each model was that each code fits:

CodeHuman applied it?JevOpenAI Decisions
Unconditional love/happy to see meYes95%100%
Affection/snugglesYes87%100%
Responsibility/need/purposeYes82%100%
ExerciseYes79%100%
Playful/friendly/entertainingNo87%100%
Personality sweet/cute/happyNo75%94%
EverythingNo87%11%
Companion/friend/familyNo59%35%
LoyaltyNo23%2%
Calm/caring/support/comfortNo23%15%
Security/protectionNo3%0%
  • Jev default (50%) applies every code it’s at least 50% sure of: 8 codes, double the human’s 4.
  • Jev adjusted to 60% drops “Companion” (59%), leaving 7 codes.
  • Decisions applies 6 codes either way. Its adjusted setting was 80%.

Both models found all four of the human’s codes and added a few more. Some extras, like “Playful”, are arguably fair readings.

How the accuracy score works

The accuracy score (F1) asks two questions, and only rewards a model that does well on both:

  • Did it add codes that don’t belong? Of the codes the model applied, what share did the human also apply? Researchers call this precision.
  • Did it miss codes? Of the codes the human applied, what share did the model find? Researchers call this recall.

The F1 score is calculated using the formula:

F1 = (2 × precision × recall) ÷ (precision + recall)

where:

precision = true positives ÷ (true positives + false positives)

recall = true positives ÷ (true positives + false negatives)

In the example above, Jev applied 7 codes, including all 4 of the human’s codes. So 57% of its codes were right, and it found 100% of the human’s. Those combine into an F1 accuracy score of 73 for that response. Decisions applied 6 codes, including all 4, for an F1 score of 80.

The scores in the rest of this article combine every code decision across all test responses.

Which is more accurate: Jev or OpenAI Decisions?

OpenAI Decisions is slightly more accurate, but only after adjustment of the strictness settings. With default settings, the two models were tied. Adjusted, Decisions was 1.7 points ahead.

Bar chart of Jev and OpenAI Decisions coding accuracy with default and adjusted strictness settings
  • The out-of-the-box gap is insignificant. A 0.7-point difference is within the margin of error.
  • The adjusted gap is small but real. A statistical check showed Decisions’ 1.7-point lead is unlikely to be chance.
  • All 2,617 responses tell the same story. Jev scored 71.4 with default settings and 74.0 adjusted. Decisions scored 72.3 and 75.4.
  • Each model has its own strength. Jev was slightly better at ranking which codes are most likely to fit. Decisions was better at the final yes-or-no call.

Coding accuracy by question: the winner changes

The overall result hides a split. After adjustment, Decisions won on pets, Jev won on urgent care, and they tied on AI.

Bar chart of Jev vs. OpenAI Decisions coding accuracy by survey question
  • Pets: Decisions led by about 4 to 5 points on both codebooks. These answers often mix overlapping ideas such as love, affection and companionship.
  • Urgent care: Jev was level on the first codebook and almost 5 points ahead on the second.
  • AI (one word): the models were within a point of each other. Both struggled with the first AI codebook, which splits single words into 13 closely related codes.

How do Jev and OpenAI Decisions compare with human coders?

Human coders don’t agree all the time either, so their agreement this time shows how hard each question is to code. On this comparison, Jev came close to the humans on urgent care, while both models fell well short on pets.

The two codebooks for each question share some codes that mean the same thing. On those codes, we measured how well the two human coders agreed. Then we let each model stand in for one coder and compared it with the other, the same way.

Bar chart of Jev and OpenAI Decisions agreement with a second human coder, against human-to-human agreement
  • Urgent care is hard even for people. The two humans agreed at only 75.8. Jev, at 74.8, came close. Decisions scored 68.6.
  • Pets is easier for people. The humans agreed at 93.6. Decisions came closer than Jev, but both were well below the humans.

Why the 50% default hurts accuracy

Adjusting the strictness setting improved each model’s accuracy by 3 to 4 points, more than the gap between the two models. At 50%, both models were too generous.

SettingModelShare of its codes the human also appliedShare of the human’s codes it found
Default strictness (50%)Jev65%81%
Default strictness (50%)OpenAI Decisions65%82%
Adjusted (60%)Jev76%75%
Adjusted (80%)OpenAI Decisions79%76%
  • Too many codes at 50%. About a third of the codes each model applied didn’t match the human’s.
  • The best setting was stricter. It ranged from 55% to 85% for Jev and from 70% to 90% for Decisions, depending on the question.
  • Decisions was the looser of the two. On urgent care at 50%, more than half of the codes it applied didn’t match the human’s codes.

The lesson learned applies to any task you use these models for: don’t trust the out-of-the-box setting. It’s necessary to adjust it using a few hundred examples a person has already coded.

Jev vs. OpenAI Decisions: cost

Both models are inexpensive, but OpenAI Decisions cost about 7 times as much as Jev in our test. For both, each code is a separate question, so the bill grows with the number of codes.

Jev or OpenAI Decisions: which should you choose?

  • Accuracy is close. Tied with default settings, and 1.7 points apart once adjusted. For many tasks, either will do.
  • Cost may decide it. Jev costs only about one seventh of OpenAI’s Decisions. At high volume, that adds up.
  • The setting matters more than the model. Adjusting the strictness setting added more accuracy than switching models.
  • Test on your own task. The winner flipped from one question to the next. A small sample of your own labeled data will tell you more than any benchmark. Our guide to ensuring accuracy in AI open-end coding shows how to check.
  • For survey coding, use a purpose-built tool. Decision models are built for fast, simple calls at scale. A survey coding platform like BTInsights works from your whole codebook and assigns codes directly.

FAQs

Is the OpenAI Decisions API more accurate than Jev?

Slightly, but only after adjustment. In our test on 2,617 human-coded survey responses, OpenAI Decisions scored 77.1 out of 100 and Jev 75.4 once each model’s strictness setting was adjusted. With default settings they were tied (72.9 vs. 72.2). Jev was more accurate on one of the three questions.

What is TypeSafe’s Jev?

Jev is a decision model from TypeSafe. It answers typed questions about a piece of text (yes/no, pick one, or a rating) and returns a confidence score instead of written text. It is built to be fast and inexpensive.

What is the OpenAI Decisions API?

The Decisions API is OpenAI’s service for fast, structured decisions. It is in public beta and runs on the gpt-6-luna model. It answers yes/no, pick-one and rating questions with confidence scores. OpenAI’s suggested uses include classifying content, routing requests and prioritizing work.

How do you measure AI coding accuracy?

Compare the AI’s codes with a human coder’s codes on the same responses. A common measure is the F1 score, shown here out of 100. It balances avoiding wrong codes (precision) with avoiding missed codes (recall). A score of 100 means the codes match exactly.

What is the cutoff, or strictness setting, in Jev and OpenAI Decisions?

Both models say how confident they are that an answer is “yes". The cutoff is how confident they must be before the answer counts as “yes”. The default is 50%, but in our test both models were more accurate with a stricter setting, between 55% and 90%.

Which is cheaper, Jev or OpenAI Decisions?

Jev. OpenAI Decisions cost about 7 times as much in our test.

Code open-ended responses in minutes

Code, review, and export open-ended responses in one workflow.

Book a Demo