Silicon Valleys Journal
  • Topics
    • Finance & Investments
      • Angel Investing
      • Financial Planning
      • Fundraising
      • IPO Watch
      • Market Opinion
      • Mergers & Acquisitions
      • Portfolio Strategies
      • Private Markets
      • Public Markets
      • Startups
      • VC & PE
    • Leadership & Perspective
      • Boardroom & Governance
      • C-Suite Perspective
      • Career Advice
      • Events & Conferences
      • Founder Stories
      • Future of Silicon Valley
      • Incubators & Accelerators
      • Innovation Spotlight
      • Investor Voices
      • Leadership Vision
      • Policy & Regulation
      • Strategic Partnerships
    • Technology & Industry
      • AI
      • Big Tech
      • Blockchain
      • Case Studies
      • Cloud Computing
      • Consumer Tech
      • Cybersecurity
      • Enterprise Tech
      • Fintech
      • Greentech & Sustainability
      • Hardware
      • Healthtech
      • Innovation & Breakthroughs
      • Interviews
      • Machine Learning
      • Product Launches
      • Research & Development
      • Robotics
      • SaaS
  • Media Kit
No Result
View All Result
  • Topics
    • Finance & Investments
      • Angel Investing
      • Financial Planning
      • Fundraising
      • IPO Watch
      • Market Opinion
      • Mergers & Acquisitions
      • Portfolio Strategies
      • Private Markets
      • Public Markets
      • Startups
      • VC & PE
    • Leadership & Perspective
      • Boardroom & Governance
      • C-Suite Perspective
      • Career Advice
      • Events & Conferences
      • Founder Stories
      • Future of Silicon Valley
      • Incubators & Accelerators
      • Innovation Spotlight
      • Investor Voices
      • Leadership Vision
      • Policy & Regulation
      • Strategic Partnerships
    • Technology & Industry
      • AI
      • Big Tech
      • Blockchain
      • Case Studies
      • Cloud Computing
      • Consumer Tech
      • Cybersecurity
      • Enterprise Tech
      • Fintech
      • Greentech & Sustainability
      • Hardware
      • Healthtech
      • Innovation & Breakthroughs
      • Interviews
      • Machine Learning
      • Product Launches
      • Research & Development
      • Robotics
      • SaaS
  • Media Kit
No Result
View All Result
Silicon Valleys Journal
No Result
View All Result
Home Technology & Industry AI

Beyond pass/fail testing: How to build reliable AI agents

SVJ Writing Staff by SVJ Writing Staff
August 5, 2026
in AI
0
Beyond pass/fail testing: How to build reliable AI agents

Software vocabulary can be extremely confusing – the difference between functional and non-functional testing is a great example. The phrase “non-functional test” suggests a test for things that don’t really matter, which is the opposite of what it actually means. However, the distinction between functional and non-functional testing is critical in software engineering. It becomes even more important with AI agents, where the failure modes are unfamiliar, and traditional testing procedures are wholly inadequate.

The distinction between “whether” and “how well”

Functional testing asks whether the software does the thing it was designed to do. When you press the button, is the form submitted? When you enter the password, do you get logged in? If you move money from account A to account B, does the right amount arrive? Each test compares an input to the output and reports the result.

Non-functional testing asks how well the software does the thing it was designed to do. How fast, how reliably, how securely, under how much load, on how many browsers, for how many simultaneous users, and with how much memory? How well does the software perform when the network is patchy, when the server restarts, and when an attacker probes the system? These are properties of the system’s behaviour rather than the behaviour itself.

Consider this analogy from the world of restaurants. A functional test investigates whether the kitchen sent out the steak for a customer who ordered steak. A non-functional test investigates how the steak arrived. Was it served while the customer was still hungry, on a clean plate, and at the right temperature? Did the kitchen also feed forty other diners that night without anyone waiting for an hour?

The history of an awkward name

The phrase “non-functional requirement” first appeared in Yeh and Zave’s 1980 paper “Specifying software requirements” in the Proceedings of the IEEE. It became more widely used after Mylopoulos, Chung and Nixon’s 1992 paper “Representing and using nonfunctional requirements”appeared in IEEE Transactions on Software Engineering. It finally reached the mainstream with Chung et al’s textbook of the same title in 2000.

There have long been engineers who disliked the term. Mike Cohn of Mountain Goat Software prefers the term “constraints” on the grounds that calling something non-functional suggests that we should not care about it. Others use “quality attributes”, and some refer to “the -ilities”, by which they mean things like reliability, scalability, usability, maintainability, portability, and observability. So the vocabulary can vary, but for nearly half a century, developers have acknowledged that the underlying distinction between what it does and how it does it is important.

Where the distinction gets fuzzy

The split between functional and non-functional testing is cleaner in theory than in practice. A login screen that takes ninety seconds to respond is functionally working, but in practical terms, it is broken. A search that returns results in the wrong order has produced an output, so strictly speaking, it passes a functional test, but again, in practical terms, the system has failed the user. Performance and correctness blur into each other at the edges, and the question of which bucket a particular test belongs in can become somewhat academic.

Applying the distinction to AI agents

An AI agent presented with a task – to book a flight, to summarise a document, to file a support ticket, to run an SQL query – can fail in two very different ways.

Functional failure is the familiar kind. The agent books the wrong flight, misreads the document, or calls the wrong tool. It returns the right answer to the wrong question. These are the agentic equivalents of a button that fails by submitting the wrong form, and they can, in principle, be tested for by comparing input-output pairs.

Non-functional failure is where agents differ from traditional software. The agent might succeed with the task on Tuesday and fail on Wednesday, despite relying on the same input, because the underlying model is stochastic. It might perform a task at ten times the budgeted cost. It might complete the task while leaking customer data into a third-party API, or be brittle when faced with adversarial inputs that no functional test would think to try.

The list of -ilities is longer for agents than for conventional software, and several items on it are genuinely new: consistency across runs, robustness to prompt injection, calibration of confidence, faithfulness to source documents, refusal behaviour, fallback when tools fail, behaviour under distributional shift, cost per successful task. Some of these have no direct analogue in conventional software, because conventional software is deterministic and does not have opinions. Agents are stochastic, and they have something close to opinions, which gives many more dimensions to the question “how well does it behave?”.

The upshot is that an organisation putting agents into production cannot rely on functional testing alone, even very thorough functional testing. A test suite which confirms that the agent got the right answer on a thousand curated cases tells you nothing about what happens on another ten thousand cases that the tests did not anticipate, or about whether the agent will pass the same tests next month, or about what the agent does when the API it depends on returns garbage. For agents, non-functional questions are not optional extras. These are questions which determine whether an agent can and should be put into production.

Functional testing asks whether software does its job. Non-functional testing asks whether anyone would want to use it when it does. The terminology is awkward, and engineers have been complaining about it since at least the early 1980s, but the underlying distinction has proved its worth. With the arrival of AI agents, it provides a useful way of thinking about what is missing from most current approaches to testing.

Previous Post

Human Decision Intelligence Is the Next Enterprise Software Category

Next Post

The End of Guesswork as a Business Model

SVJ Writing Staff

SVJ Writing Staff

Next Post
The End of Guesswork as a Business Model

The End of Guesswork as a Business Model

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

  • Trending
  • Comments
  • Latest
Faith and the Digital Transformation of Religion: How One Person Began Helping Faith Communities and People of Faith

Faith and the Digital Transformation of Religion: How One Person Began Helping Faith Communities and People of Faith

December 30, 2025
The AI Cold War and How to Prepare for It

The AI Cold War and How to Prepare for It

May 1, 2026
AI’s Most Underrated Role: Giving Enterprise Architects Back Their Focus

AI’s Most Underrated Role: Giving Enterprise Architects Back Their Focus

November 26, 2025
The UK’s Seed-to-Series A gap is growing. Should we fix it?

The UK’s Seed-to-Series A gap is growing. Should we fix it?

November 25, 2025
The Human-AI Collaboration Model: How Leaders Can Embrace AI to Reshape Work, Not Replace Workers

The Human-AI Collaboration Model: How Leaders Can Embrace AI to Reshape Work, Not Replace Workers

1

50 Key Stats on Finance Startups in 2025: Funding, Valuation Multiples, Naming Trends & Domain Patterns

0
CelerData Opens StarOS, Debuts StarRocks 4.0 at First Global StarRocks Summit

CelerData Opens StarOS, Debuts StarRocks 4.0 at First Global StarRocks Summit

0
Clarity Is the New Cyber Superpower

Clarity Is the New Cyber Superpower

0
The Intelligence Company Hiding Inside an AI Narrative

The Intelligence Company Hiding Inside an AI Narrative

August 5, 2026
The Central Brain: How Human Behaviour Becomes Software

The Central Brain: How Human Behaviour Becomes Software

August 5, 2026
The Market Field Intelligence Layer Investors Are Missing

The Market Field Intelligence Layer Investors Are Missing

August 5, 2026
The End of Guesswork as a Business Model

The End of Guesswork as a Business Model

August 5, 2026

Recent News

The Intelligence Company Hiding Inside an AI Narrative

The Intelligence Company Hiding Inside an AI Narrative

August 5, 2026
The Central Brain: How Human Behaviour Becomes Software

The Central Brain: How Human Behaviour Becomes Software

August 5, 2026
The Market Field Intelligence Layer Investors Are Missing

The Market Field Intelligence Layer Investors Are Missing

August 5, 2026
The End of Guesswork as a Business Model

The End of Guesswork as a Business Model

August 5, 2026

About & Contact

  • About Us
  • Branding Style Guide
  • Contact Us
  • Help Centre
  • Media Kit
  • Site Map

Explore Content

  • Events
  • Newsletter
  • Press Releases
  • Reports & Guides
  • Topics

Legal & Privacy

  • Advertiser & Partner Policy
  • Communications & Newsletter Policy
  • Contributor Agreement
  • Copyright Policy
  • Privacy Policy
  • Prohibited Content Policy
  • Terms of Service

Tiny Media Brands

  • Silicon Valleys Journal
  • The AI Journal
  • The City Banker
  • The Wall Street Banker
  • World Lifestyler
  • About
  • Privacy & Policy
  • Contact

© 2025 Silicon Valleys Journal.

No Result
View All Result

© 2025 Silicon Valleys Journal.