The list that keeps getting longer
“Find every company that fits these criteria.”
It sounds like a simple assignment until the criteria arrive. The companies must operate in certain countries, sell a particular product, have announced funding recently, and provide evidence for every entry. Finding names is only the start. Someone must check each one, fill in the missing details, remove duplicates and attach sources.
On 25 September 2026, Exa launched Agent Ultra, the highest-effort mode of its existing AI research agent. Exa says it can break a large assignment into smaller searches, coordinate agents across many sources and assemble the findings into a usable result. It is available through the Exa Agent API.
The pitch is appealing: hand over the messy research brief and receive something closer to a checked spreadsheet than a pile of search results.
There is a catch. More effort means more time and potentially more cost. Exa’s own documentation says Ultra typically takes about 30 minutes on complex tasks and can take up to three hours on particularly challenging ones. This is a tool for work you can send away and revisit, not a quick answer while your coffee is still hot.
An upgrade to an existing agent
Exa first introduced Exa Agent in June 2026. That product already combined web search, page reading and reasoning for tasks such as company research, list building and data enrichment. Ultra expands how much work the agent can put into a difficult assignment; it is not Exa’s first research agent.
The distinction helps explain the new release. A conventional search API returns pages for an application or person to examine. Exa Agent can take on a larger workflow: search from different angles, read candidate pages, check entries against a brief and assemble an answer. Ultra tells that agent to use its most intensive mode.
Exa positions it for jobs where missing half the relevant results would defeat the exercise. Think of a market map, a detailed literature search or a list of companies meeting several narrow requirements. A fast answer with five plausible entries might look impressive. It may also be useless if the assignment called for fifty.
Exa’s Agent documentation says a run can return prose, structured data and citations tied to its findings. Developers can build those results into their own software rather than manually copying material out of a chat window.
How the research spreads out
A big search task rarely follows one straight line. Consider a request to find battery-recycling companies across Europe, then add each company’s founders, funding and customers. One query might uncover a company. Another might confirm what it actually does. Several more could be needed to establish the requested details.
Exa says its agent splits such work into subtasks and assigns agents to pursue different parts at once. It can use more capable models where the task needs them and faster models where they are sufficient. Ultra gives that process a larger work budget so it can keep looking for qualified results.
That is the practical meaning of “higher effort.” It does not make the system automatically correct. It gives the system room to search more, investigate more candidates and follow leads that a cheaper run might leave unexplored.
Exa also lets a developer supply an existing list and ask Ultra to find additional entries while excluding those already collected. That matters for recurring work. A research team may have twenty verified companies today and want twenty more next week. Starting from zero each time would waste effort—and invite duplicates.
The useful output is more than an answer
A confident paragraph can hide a disappointing research job. “Here are the leading companies” sounds complete even when it names only the easiest ones to find.
Structured output makes omissions easier to spot. A developer can specify fields such as company name, official website, funding round and evidence link. Exa says its Agent API can return data in a requested JSON structure, alongside grounding information and a cost breakdown. In ordinary language, the result can resemble a table with a place for supporting sources.
That structure helps downstream software. A company could send verified candidates to an internal review queue, compare them with an existing database or flag rows whose sources need human inspection.
Still, a correctly shaped record is not necessarily a true one. A field labeled “customer” may contain a partner, an old customer or a claim repeated across several websites without independent confirmation. The value of a citation depends on what its page actually says.
Exa’s documentation says completed runs can include citations for text or individual structured fields when emitted. That wording matters. Teams using Ultra for important decisions should inspect the returned evidence and decide what to do with fields the system could not establish.
A benchmark built for sprawling assignments
To support its launch claims, Exa points prominently to WANDR, a research-agent benchmark. WANDR tests whether an agent can find a broad set of qualifying items and investigate each one deeply enough to support its claims. The benchmark paper, published in August 2026, describes 500 challenging data-collection tasks.
A task may require far more than a name and a link. It can ask for entities, their relationships, supporting excerpts and a target number of completed records. That makes WANDR a closer fit for a market map than a trivia test.
The benchmark attempts to check evidence as well as answers. Its method retrieves cited pages and evaluates whether they support the submitted records. It reports measures of coverage and accuracy, including scores that give partial credit and stricter scores that require a complete match.
Why bother with all that machinery? Because an agent can fail in several ways. It may overlook a valid company, identify one but miss its funding details, confuse two similarly named firms, or cite a page that does not prove the claim. “Got an answer” tells us very little about those failures. WANDR tries to make them visible.
What Exa says Ultra achieved

In its launch post, Exa says Agent Ultra performed better than the compared maximum-effort settings of Claude Opus 5.5, GPT-6 Astra and Perplexity Agent across several research benchmarks. It reports a 12.6% advantage over Opus 5.5 on its WANDR comparison, while claiming lower average cost per task. It also reports gains on DeepSearchQA and WideSearch. These are Exa’s reported evaluations, not independent tests of the new release.
The company’s largest percentage claims concern its own Find-All Company test. Exa says Ultra found dramatically more qualifying entities than the systems it compared. That is interesting for list-building customers, but readers should treat a company-designed test differently from a broadly reproduced public benchmark.
Exa provides a useful methodological detail about its WANDR run. It says its grader shares evaluation logic with the published benchmark, while differing in the contents tool, transport logic and judge model. For competitor results, Exa says it used previously published scores where they came from that grader setup and ran tests itself where they did not.
Those details make the claim more inspectable. They do not remove every comparison problem. Tool access, settings, budgets and evaluation choices can affect what an agent finds.
Even the strongest agents miss things
The WANDR paper offers a useful reality check for any victory lap. In the systems evaluated for that paper, the strongest high-effort result reached 0.363 soft F1 and 0.133 hard F1. Those figures came from the benchmark research, not from Exa Ultra’s September launch test. They show how demanding wide-and-deep research remains.
You do not need to decode the scoring formula to grasp the point. Finding many entries, checking every detail and supplying adequate evidence is hard. Performance drops as the required list grows and each item needs more layers of investigation.
The paper identifies several failure points: agents can miss candidates, fail to enrich the ones they find or provide incomplete supporting evidence. Some pages are also difficult to retrieve because of broken links, login requirements or other access barriers.
That context makes Ultra’s direction more meaningful. Exa is spending more effort on the parts of research that short answers often gloss over: breadth, qualification and proof. It also sets a sensible expectation. A stronger research agent can reduce manual work while still needing review, especially when a missed company or an unsupported claim could change a business decision.
Where a longer search could pay off
Exa offers several examples of work it thinks suits Ultra. An AI company might seek papers and code repositories that implement a particular technique, then check whether each result has a reproducible evaluation setup. A financial research team might build a map of companies in a narrow industry. A sales operations team might look for businesses that meet detailed product and hiring criteria.
Those examples share a pattern: the deliverable is a set of qualified records, not one clever sentence. Completeness matters because the missing entries may be exactly the ones the researcher hoped to discover.
The value of extra search effort will vary by task. Spending half an hour to confirm a fact that appears on an official homepage makes little sense. Spending that time to investigate an unfamiliar market may be reasonable. A team can also divide the job: let Ultra gather a candidate list, then have a person review the entries most likely to affect a decision.
That final review should follow the evidence, not just the polished summary. A cited company page may confirm a product but say nothing about the funding round in the same row. Useful research makes it possible to catch that gap before the row travels into a report.
The clock runs longer on purpose
Ultra is an asynchronous workflow. A developer starts a run, saves its identifier, then checks or streams its progress until it finishes. Exa says a complex run typically completes in about 30 minutes, while very challenging work can take up to three hours.
That makes it a poor fit for a chatbot expected to answer a customer before they leave the page. It makes more sense for an analyst preparing tomorrow’s briefing or a software system refreshing a detailed research list in the background.
Exa documents controls for both spending and duration. Ultra’s default maximum cost is $20 per run, and developers can set their own cost cap within the documented range. They can also set a time limit or stop a run early. When a budget limit approaches, the agent stops starting new work and returns what it has found.
That last phrase deserves attention. A returned result after a limit is reached may be useful, but it may be incomplete. Exa includes a stopReason so applications can distinguish a finished task from a run that stopped because it exhausted its budget or time.
More effort has a price
Ultra does not have a simple flat fee of $1 per request. Exa’s pricing page lists $1 for its fixed x-high effort setting, but its Ultra documentation describes Ultra as metered by usage, with the default $20-per-run cap. Those are different settings.
For metered Agent work, Exa lists $0.10 per Agent Compute Unit and $0.005 per search tool call. Email and phone contact enrichment can add separate charges. A run that finishes early can cost less than its cap; a task that fans out across many searches and reasoning steps can cost more than a modest lookup.
The relevant business question is not “Is one run cheap?” It is “What does a usable, checked result cost?” A $10 run that produces an incomplete list may create hours of follow-up work. A more expensive run could save time if it supplies better coverage and evidence. The reverse can also happen.
That is why a pilot should use real briefs. Compare Ultra with a lower-effort setting on the same assignment. Count qualified entries, inspect their citations, note how many need correction and record the full cost. The cheapest route becomes clearer once the checking work joins the bill.
What comes after the first result
Research rarely ends when the first table arrives. A teammate may ask for more companies, narrower criteria or a closer look at one candidate. Exa says Agent runs can continue from previous work, and Ultra can accept existing rows so it looks for additional results rather than repeating the same list.
That could make the agent useful as a research partner over several passes. Start broad. Inspect the candidates. Tighten the rules. Ask for more evidence where the sources look thin.
The process still needs a clear brief. If “AI infrastructure company” means one thing to a venture investor and another to a cloud engineer, the agent cannot quietly resolve the disagreement for both. Explicit criteria make its results easier to judge. So do requirements for official sources, recent dates or an honest “cannot verify” where evidence is absent.
An independent roundup from The Neuron describes the launch as aimed at extensive list-building and similarly identifies the benchmark comparisons as vendor claims. That is the right starting posture for users: the product is available to try, and its strongest performance promises need testing against work that matters to them.
The real test is the missing row

Agent Ultra’s most interesting promise is not that it can produce a longer answer. It is that it might find the qualified company, paper or person a shorter research pass would miss—and show enough evidence for someone else to check the discovery.
Exa has launched a tool designed to spend more time and compute pursuing that goal. Its API supports structured results, citations, progress tracking and spending limits. Its launch benchmarks suggest strong performance, but Exa ran the new comparisons and published the claims itself. The independent WANDR research also makes clear how much room every research agent has to improve.
For developers and research teams, the sensible experiment is straightforward. Give Ultra a difficult assignment with a known review process. Check the sources. Look for duplicates and omissions. Compare the finished work with a lower-effort run and with the time a human team would have spent.
If Ultra consistently uncovers well-supported entries that other approaches overlook, a half-hour wait could be a bargain. If it merely produces a bigger list that still needs extensive repair, the impressive swarm of agents will have left the hardest part on a person’s desk.
Sources
- Exa, “Introducing Exa Agent Ultra,” 25 September 2026
- Exa, Agent Ultra documentation
- Exa, Agent API documentation
- Exa, API pricing
- Exa, original Exa Agent announcement, 16 June 2026
- WANDR benchmark paper, 14 August 2026
- The Neuron, AI news roundup covering Agent Ultra, 26 September 2026
The Kingy Brief
Get the next Kingy Brief.
Source-checked AI changes, original tests and one practical thing to try.
Free · Choose your subjects · Double opt-in · Unsubscribe anytime
