All guides
AI in B2B Sales 11 min read

Which LLM Is Best for Automated Company Research? Our Benchmark of 18 Models

We compared 18 language models on 100 anonymised research cases: answer quality, cost per 1,000 companies, model time and valid JSON. Plus two lessons: well-known companies measure memory, and the harness matters as much as the model.

LC
Founder, CegTec · 9 October 2026

Which LLM is best for automated company research?

We tested 18 language models on the same 100 anonymised research cases. Six models are within the margin of error on answer quality: claude-haiku-5.5, deepseek-v4.1-flash, gemini-3.5-flash-lite, gpt-6-luna, qwen3.5-flash and our reference gemini-3.1-flash-lite. Among them, cost, runtime and whether the model runs cleanly in our agent decide. deepseek-v4.1-flash cost a little over a third of the reference at the same quality. The test also showed that a replay result is not enough: in the full agent run, part of the field did not run at all.

The measurements are ours, from developing the Research Agent. We sell a research product and have an interest in the answer. That is why the method is below, and why it also says what we did not measure.

Results at a glance

Replay on 100 cases, sorted by answer quality. Each case runs a page filter (is this page relevant?) on all recorded pages and an answer step (task plus pages into JSON by schema). Cost is the token list price, scaled to 1,000 cases. Model time is the sum of the model calls per case, not the time of a whole research run.

ModelAnswer quality (1 to 5)Factual accuracyCost per 1,000 casesModel time per caseValid JSON (raw, out of 100)
claude-haiku-5.54.004.71$0.611.9 s98
deepseek-v4.1-flash3.994.55$0.293.8 s100
gemini-3.5-flash-lite3.974.48$1.112.0 s100
gpt-6-luna3.884.60$0.364.1 s100
qwen3.5-flash3.864.12$0.212.1 s100
gemini-3.1-flash-lite (reference)3.853.99$0.882.4 s100
gpt-5-mini3.714.08$0.945.0 s100
glm-5.3-flash3.704.31$0.3712.0 s98
qwen3.8-flash3.663.88$0.454.0 s100
gemma-4-31b-it3.483.99$0.4116.8 s100
qwen3.6-35b-a3b3.392.87$0.352.9 s92
qwen3.8-27b (reasoning off)3.293.76$0.493.2 s100
nemotron-3.5-lightning3.102.14$0.193.7 s100
gemma-4-26b-a4b-it3.083.63$0.224.4 s98
qwen3.8-27b (reasoning low)3.073.68$2.2613.5 s100
gpt-5-nano3.063.24$0.205.8 s100
deepseek-v3.22.882.79$0.8142.3 s98
mimo-v2.6-flash2.773.33$0.583.3 s70
deepseek-v4-flash2.742.91$0.308.9 s83

As of: runs of 7 October 2026. Each model ran against one fixed endpoint, with its own reasoning and output-format setting. Failed calls count as invalid. The quality score is the mean of factual accuracy, completeness and format.

How to read the table:

  • The first six rows cannot be separated. Quality per case varies with a standard deviation of 0.75, and the mean over 100 cases is accurate to about ±0.15. Against the reference, deepseek-v4.1-flash was +0.14 and gemini-3.5-flash-lite +0.12, both at the edge of the margin of error.
  • Completeness is low for everyone, between 1.9 and 2.6, because the replay knows only the pages that were already recorded. The differences in the overall score come mainly from factual accuracy.
  • Cost varies more than quality. Between the reference ($0.88) and deepseek-v4.1-flash ($0.29) there is a factor of about three at the same score. Reasoning makes things dearer: qwen3.8-27b with reasoning on low cost 4.6 times as much and scored lower (3.07 instead of 3.29).
  • Valid JSON is not a given. mimo-v2.6-flash returned valid JSON in only 70 of 100 cases, deepseek-v4-flash in 83. Some answers ran in loops up to the token limit.
  • Schema-valid does not mean useful. gpt-5-nano was 100% valid but padded its answers with placeholders: 30 entries of the “Not found” type, 2.45 leads per answer against 1.33 for the reference.

The test in the full agent run

The replay tests only the answer step. So three selected models also ran through the complete agent loop: search, page fetch, page filter, several steps, output by schema. Ten companies, standard depth, no fallback to another model. A run counted as usable if the agent finished on its own, the JSON was schema-valid and the chosen model served every call. We did not rate answer quality in this test.

ModelUsable runsSeconds per run (median)Cost per run (mean)Cost vs reference
gemini-3.1-flash-lite (reference)10 of 1069$0.00421.00
deepseek-v4.1-flash10 of 1017$0.000950.23
qwen3.5-flash, default setting0 of 1090$0.01305.1
qwen3.5-flash, reasoning off (test only)9 of 1015$0.00231.2
gpt-6-luna, default setting2 of 10122$0.00903.6
gpt-6-luna, reasoning off (test only)5 of 1052$0.00571.9

The qwen and luna rows come from an earlier run the same day, with a reference on the first five companies (median 16 s, $0.0028 per run). The reference was slower in the later run and searched more, 4.8 times per run instead of 2.6. Against the earlier run, deepseek-v4.1-flash is at 0.31 of the cost. That fits the roughly 33% from the replay.

In the replay, qwen3.5-flash and gpt-6-luna were on a par with the reference. In the agent run, on default settings, they barely got through.

What we learned

The harness matters as much as the model

Our agent works in steps: thought, action, observation, exactly one step per reply. Models that write several steps into one reply are rejected, and the searches then never run. That affected qwen3.5-flash and gpt-6-luna in the agent run. It also affected deepseek-v4.1-flash, mainly when building company lists: 2 of 3 lists were produced, at 172 to 602 seconds per list and 24 rejected steps. The rows that came out were real, matching companies. The problem sits in the format, not in the model’s knowledge or judgement.

On 9 October 2026 a DeepSeek run in the live app returned “Not found” after about five seconds. It was the same mechanism, not a verdict on the model. The same model was among the best in the replay. Anyone who judges a model only on the answer step misses whether it gets to answer at all inside their own agent. We are adjusting the harness and will then measure again with the same lists and companies.

Well-known companies measure memory, not research

Our live tests ran on publicly well-known companies. Many models know something about them without searching. In the test of qwen3.5-flash with reasoning off, 4 of the 9 usable runs made not a single tool call: the model answered from memory. Such runs count as usable in the test and say little about research on an unknown mid-sized company, where there is nothing to recall from training. A reliable live test needs companies a model cannot know anything about.

Eval and production measure different things

In July we ran two full agent-loop runs, each with 12 tasks, live search and quick research depth. All models tested were 100% reliable and schema-valid. Quality in July was 3.44 to 3.56 (first run) and 2.94 to 3.61 (second run). Same models, two runs, different values: gemini-3.1-flash-lite cost $0.00244 per run in the first and $0.00292 in the second, gpt-5-nano $0.00061 and $0.00094.

July 2026, agent loop, 12 tasksQuality (judge)Cost per run
gemini-3.1-flash-lite, run 1 / run 23.47 / 3.47$0.00244 / $0.00292
gpt-5-nano, run 1 / run 23.56 / 3.58$0.00061 / $0.00094
deepseek-v3.2, run 13.50$0.00056
deepseek-v4-flash, run 13.44$0.00179
gpt-5-mini, run 13.56$0.00079
gemini-2.5-flash, run 23.61$0.01030
gemini-2.5-flash-lite, run 22.94$0.00294
gemini-3.5-flash, run 23.42$0.03134

deepseek-v3.2 reached 3.50 in July in the agent run with n=12 and 2.88 in October in the replay with n=100, with 42 seconds of model time per case. The measurements are not comparable: different tasks, different scale, different test set-up. That is exactly the lesson. A single benchmark value for a model is not a property of the model, but of model, test set, harness and judge together.

Also in July: gemini-2.5-flash-lite cost as much as the reference at measurably lower quality (2.94 against 3.47). gemini-3.5-flash cost 10.7 times the reference with no quality advantage (3.42).

Method and limits

  • Test set: the first 100 cases of an anonymised test set of real research tasks, with no individual companies named. In July it was 12 tasks in the agent loop.
  • Replay: per case, a page filter on all recorded pages and an answer step by schema. Search is not part of the replay.
  • Judge: claude-haiku-4.5, not a candidate in the replay, rates factual accuracy, completeness and format against the source text, each from 1 to 5. For claude-haiku-5.5 we cross-checked with a judge from another vendor (mistral-medium-3.1): 4.05 against 3.84 for the reference, in the main judge 4.00 against 3.85. Both judges put the model ahead of the reference.
  • Prices: list price per million tokens (input / output) on the endpoint used, as of 7 October 2026. Billed cost matched list price with a few exceptions: gpt-6-luna adds 25% to the input price for prompts of 1,024 tokens or more, so its answer step was billed at 1.21 times list price.
  • Latency: measured is model time per call and per case, not the time of a research run from request to answer. End-to-end latency: not measured.
  • Not measured: cost per company in live production, quality of answers in the full agent run, research on unknown companies in the live test.
  • Noise: at n=100 quality is accurate to about ±0.15, at n=12 considerably less. For rankings within 0.15 points the test is too weak.

What you can take from this

  1. Base the model choice on cost and reliability, and use quality only as a minimum threshold. For us, the case for a cheaper model was not the score but a similar quality level at a third of the cost.
  2. Test the model in your own harness, not only on the answer step. Count rejected steps, runs without a tool call and runs that hit the time limit.
  3. Test with companies the model does not know. Otherwise you measure memory.
  4. Count raw JSON, not repaired JSON. A repair stage before evaluation makes results cleaner than production is.
  5. Run it twice. The same test gave us different costs for the same model.

Try it yourself

The Research Agent is our product for live research on companies and leads, via API or MCP from Claude, Cursor and other MCP clients. You can create an account at app.research-agent.net. If you want to know which set-up fits your own research needs, book a free intro call. We look at your use case and tell you honestly whether a change pays off.

Sources

All figures come from the Research Agent team’s own measurements: replay benchmark and live agent check of 7 to 9 October 2026 (18 models, 100 cases, 10 companies in the agent run), agent runs of 11 and 13 July 2026 (12 tasks). Prices follow list prices as of 7 October 2026. Model names and versions are those of the respective vendors. Features and prices change quickly. The results apply to our test set-up and are not a general ranking of the models.

LLM benchmarkCompany researchResearch AgentAI agentsModel comparison

Common questions

Which LLM is best for automated company research?

On our 100 anonymised cases, claude-haiku-5.5 (4.00), deepseek-v4.1-flash (3.99), gemini-3.5-flash-lite (3.97), gpt-6-luna (3.88), qwen3.5-flash (3.86) and our reference gemini-3.1-flash-lite (3.85) were close together on answer quality. The margin of error of the mean is about ±0.15 on a scale of 1 to 5, so a ranking within that group is not reliable. Cost and reliability separate the field more clearly: at similar quality, deepseek-v4.1-flash cost $0.29 per 1,000 cases and the reference $0.88.

How much does it cost to research a company with an LLM?

In our replay, which measures only the page filter and the answer step, models ranged from $0.19 (nemotron-3.5-lightning) to $2.26 (qwen3.8-27b with reasoning) per 1,000 cases, the reference at $0.88. That is not the price of a full research run with search and several steps. In the full agent run on ten companies, the reference cost $0.0042 per run on average and deepseek-v4.1-flash $0.00095. We have not measured cost per company in live production.

Why is an eval not enough to choose a research model?

Because an eval measures only what it measures. Our replay feeds pre-recorded pages and tests the answer step. In the full agent run three more effects appeared: the agent expects one step per reply, some models write several, and the searches then never run. qwen3.5-flash managed 0 of 10 runs that way, and 9 of 10 with reasoning switched off. Tests on very well-known companies also partly measure memory instead of research. A model should be tested in the real harness on companies it does not know.

Is DeepSeek a poor model for company research?

No. deepseek-v4.1-flash scored 3.99 in the replay, among the best, and at $0.29 per 1,000 cases was among the cheapest. When a DeepSeek run in the live app ended after about five seconds with "Not found" on 9 October 2026, the cause was a format problem in our agent harness, not the model: DeepSeek writes several steps into one reply and the agent processes only one per reply. Earlier DeepSeek variants scored clearly lower in the same replay, and that is in the table too.

How reliable are the numbers?

Cost, runtime and the share of valid JSON answers are reliable because they are counted rather than rated. Quality comes from an LLM judge (claude-haiku-4.5) on a scale of 1 to 5 and is noisy: the standard deviation per case is 0.75, and the mean over 100 cases is accurate to about ±0.15. The judge is itself a model, so for one candidate we cross-checked it with a second judge from a different vendor. End-to-end latency per company and quality in the full agent run were not measured.

Next step

GTM Goat runs exactly what you just read.

Book an intro call 30 minutes, then 4 weeks free
Rather start on your own? Pricing and free trial →

Playbooks für B2B Outbound freischalten

Kostenlos. E-Mail eintragen → Passwort erhalten → Playbooks lesen.