Test the task
A model can win one test and still lose on the job you need. Assaf Elovic shared GPT Researcher’s comparison of context filters across 28 research tasks. Jev kept a higher share of relevant passages than embeddings: 73% against 46%.
That measures the passages kept, not a 59% improvement in finished reports. The test used one writer model and one report type. Its writer and judge came from the same model family, which limits how far we’d carry the result.
Glean’s separate tests found a different answer for each task. Jev improved routing, but a tuned classifier and Glean’s existing ranker did better elsewhere. Repeatable citation judgments didn’t establish accuracy without human labels.
So our takeaway is to name the decision before picking a model. Keep the old method in the comparison, including a simple keyword filter where it fits. Count the wrong answers and the extra review they create, alongside the time and cost. Use the same records in each run so you can explain what changed.
A faster call isn’t useful if someone has to repair its answer. You should test the whole job before changing the part.
Assaf Elovic’s announcement, . GPT Researcher’s method and results, accessed September 28, 2026.
Eddie Zhou, Chau Tran (@mr_cheu), @MatZhao, Aviral Singh and Manav Agrawal. Is Jev overhyped? We tested it on 4 real enterprise tasks. Published by Tony Gentilcore on X, . Commentary added September 28, 2026.
