Close Menu
The Financial News 247The Financial News 247
  • Home
  • News
  • Business
  • Finance
  • Companies
  • Investing
  • Markets
  • Lifestyle
  • Tech
  • More
    • Opinion
    • Climate
    • Web Stories
    • Spotlight
    • Press Release
What's On

David Ellison’s $81B merger of Paramount and Warner Bros. Discovery officially completed

October 6, 2026

Nobody Can Tell You What An AI Task Costs: Not Even The AI

October 6, 2026

Context Engineering Decides Whether Your AI Testing Agents Are Useful

October 6, 2026

Wall Street on track for record-smashing $90B year: New York State comptroller

October 6, 2026

‘Jurassic World: Dominion’ Final Cost Soars To $684 Million

October 6, 2026
Facebook X (Twitter) Instagram
The Financial News 247The Financial News 247
Demo
  • Home
  • News
  • Business
  • Finance
  • Companies
  • Investing
  • Markets
  • Lifestyle
  • Tech
  • More
    • Opinion
    • Climate
    • Web Stories
    • Spotlight
    • Press Release
The Financial News 247The Financial News 247
Home » Context Engineering Decides Whether Your AI Testing Agents Are Useful

Context Engineering Decides Whether Your AI Testing Agents Are Useful

By News RoomOctober 6, 2026No Comments5 Mins Read
Facebook Twitter Pinterest LinkedIn WhatsApp Telegram Reddit Email Tumblr
Share
Facebook Twitter LinkedIn Pinterest Email

Mayank Bhola, Co-Founder and Head of Products, TestMu AI (Formerly Lambdatest).

​Every QE leader who has piloted AI test generation has some version of the same story. The agent produced an impressive volume of tests in minutes, but on closer inspection, many checked things nobody cared about, referenced UI elements that did not exist or retested the same happy path in slightly different ways.

The instinct is to blame the model. A better diagnosis is that the agent was working blind. It knows what a login page generally looks like but not what your login page does, how users depend on it or where it broke last quarter.

I have spent a decade building testing infrastructure and the last few years building testing agents, and I now believe the gap between disappointing pilots and production value has one name: context engineering. It is becoming one of the highest-leverage skills in quality engineering, and most teams have barely started practicing it.

Why Testing Agents Fail Without Context At Scale

Applied AI spent its first few years focused on prompt engineering, the craft of phrasing instructions so a model responds well. As agents began taking on multistep work, the focus shifted.

Andrej Karpathy gave the shift a clear articulation in 2025, arguing that the real work is filling the model’s context window with the right information: instructions, retrieved knowledge, memory, tool definitions and prior outputs, structured so the model can use them.

A testing agent depends heavily on knowledge that never appears in a model’s training data. Which user journeys drive revenue? Which module has a history of regressions? What did the requirement actually mean rather than what the ticket happened to say? Which failures are tolerable, and which lead to an executive escalation?

Strip that away, and the agent defaults to the statistical average of every application it has seen, producing tests for a generic app rather than yours.

The research record supports this. A 2026 study of an autonomous QA agent for Selenium script generation found that standard LLM generation achieved a 30% execution success rate. Grounding the same task in project documentation and the application’s actual DOM structure through retrieval increased success to 90% while largely eliminating hallucinated UI elements.

Same model, same task, but triple the reliability, and the variable was the context.

Context Rot: More Context Is Worse Context

The obvious response is to give the agent everything: every requirement document, every old test plan and the full company wiki. The evidence suggests that it can backfire.

Researchers describe context rot, where model performance degrades as input length grows. A Databricks study also found accuracy declining around 32,000 tokens, well below what modern context windows technically allow.

The goal is not maximum context. It is the smallest set of information that maximizes the likelihood of the right outcome.

For testing, that means curation is the real work.

An agent testing a checkout flow may need payment requirements, the DOM of the relevant pages, defect history for that flow and the top user paths through it. It does not need the entire product wiki. Adding irrelevant material can make the result worse.

The Four Moves Of Context Engineering, Translated For QE Teams

LangChain’s framework reduces context work to four operations: write, select, compress and isolate. Translated into a QE organization, they show where the real effort goes.

• Write: Turn tribal knowledge into machine-usable information: risk priorities, coverage philosophy and definitions of what constitutes a meaningful failure. Much of this knowledge sits in the heads of a handful of senior testers. Encoding it turns tacit expertise into reusable infrastructure.

• Select: Retrieve context for the journey being tested rather than preloading everything. A checkout agent should receive checkout requirements, defect history and relevant usage patterns, not unrelated product documentation.

• Compress: Summarize where the summary retains the important signal. A concise digest of last quarter’s payment regressions can be more useful than thousands of tokens of raw bug reports.

• Isolate: Keep unrelated context out. Irrelevant information does not simply sit harmlessly in the context window; it can distract the agent and degrade its decisions.

The shortcut many teams miss is that much of this context already exists in structured form.

A Jira ticket is context: intent, acceptance criteria, discussion history and links to the flows it touches. When we designed KaneAI, one of the decisions I defend most strongly is starting from artifacts teams already maintain rather than asking them to create an entirely new knowledge base for the agent.

Generating tests directly from a ticket is a simple example, but it illustrates the point. Once the agent has something specific to your application, it stops testing the average one.

How To Put Context Engineering Into Practice

Step 1: Build one context pack. Pick your highest-value user journey and collect its requirements, relevant page structures, recent defects and most important user paths. Make that material available to your testing agent.

Step 2: Run the comparison. Test the same journey with and without the context pack. Compare executable tests, hallucinated elements, duplicated scenarios and coverage of critical paths. The difference gives you a concrete business case.

Step 3: Assign ownership. Name a senior tester or quality engineer as the context owner for each critical flow and set a refresh cadence. Track failures caused by stale context separately from failures caused by product code.

Teams disappointed by AI testing agents and teams getting real value from them may use similar models. What differs is what those agents know.

The models are rented, but the context is yours.

The Bottom Line

Start with one high-value user journey and build its context pack: requirements, relevant page structures, recent defect history and the most important user paths.

Run your testing agent with and without that context and compare the results.

That difference will tell you far more about the future value of AI testing in your organization than another model benchmark ever will.​

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Mayank Bhola
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Related News

Nobody Can Tell You What An AI Task Costs: Not Even The AI

October 6, 2026

When Your Data Architecture Becomes The Constraint On Your Application

October 5, 2026

Preparing For The Internet’s Next Cryptographic Transition

October 5, 2026

​Why Every AI Strategy Needs A Workforce Strategy

October 5, 2026

Building Healthcare AI That CARES

October 5, 2026

CISOs Are Missing The Most Vulnerable Attack Vector

October 5, 2026
Add A Comment
Leave A Reply Cancel Reply

Don't Miss

Nobody Can Tell You What An AI Task Costs: Not Even The AI

Tech October 6, 2026

Konstantin Klyagin is the Founder of Redwerk and QAwerk, driving innovation in custom software development…

Context Engineering Decides Whether Your AI Testing Agents Are Useful

October 6, 2026

Wall Street on track for record-smashing $90B year: New York State comptroller

October 6, 2026

‘Jurassic World: Dominion’ Final Cost Soars To $684 Million

October 6, 2026
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
Our Picks

Whole Foods’ private-label ‘365’ chicken comes from Perdue plants with salmonella issues: report

October 5, 2026

‘I didn’t know what that was until about a week ago’

October 5, 2026

Hooters shutters last Jacksonville, Florida, location after bankruptcy restructure

October 5, 2026

Las Vegas outpost of The Corner Store coming this fall

October 5, 2026
The Financial News 247
Facebook X (Twitter) Instagram Pinterest
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact us
© 2026 The Financial 247. All Rights Reserved.

Type above and press Enter to search. Press Esc to cancel.