Penetration Testing

Test OWASP A10:2025 Mishandling of Exceptions in 2026

Tejas K. Dhokane, Marketing Associate at AppSecure Security
Tejas K. Dhokane
Marketing Associate
A black and white photo of a calendar.
Updated:
September 11, 2026
•
A black and white photo of a clock.
12
mins read
Tejas K. Dhokane, Marketing Associate at AppSecure SecurityVijaysimha Reddy, Security Engineering Manager at AppSecure
Written by
Tejas K. Dhokane
, Reviewed by
Vijaysimha Reddy
A black and white photo of a calendar.
Updated:
September 11, 2026
•
A black and white photo of a clock.
12
mins read
How to Test OWASP A10:2025 Mishandling of Exceptional Conditions
On this page
Share

Testing OWASP A10:2025 Mishandling of Exceptional Conditions means forcing an application to fail — through malformed input, dependency timeouts, resource exhaustion, or unexpected state transitions — and verifying that the failure never leaks data, crashes into an insecure default, or opens a bypass path. This guide walks through the manual test methodology used to validate exception handling across APIs, business logic, cloud infrastructure, mobile clients, and AI pipelines in 2026.

TL;DR

  • Exception handling failures surface through forced errors: malformed input, dependency timeouts, and resource exhaustion reveal stack traces and fail-open logic.
  • Manual testing finds broken exception paths that automated scanners miss because they require multi-step state manipulation.
  • PCI DSS, SOC 2, and ISO 27001 all expect evidence that error conditions don't disclose sensitive data or bypass controls.
  • AppSecure Security tests exception handling manually across APIs, mobile apps, and AI pipelines as part of every penetration testing engagement in 2026.

Why This Matters

Unhandled exceptions are rarely dramatic. A single malformed API request, a timed-out payment gateway call, or a database connection that drops mid-transaction produces a stack trace, a 500 error, or a silent fallback to a less secure state. Attackers use these moments deliberately — forcing an exception is a reconnaissance technique, not an accident, and the response tells them exactly how the system behaves under stress.

OWASP A10:2025 Mishandling of Exceptional Conditions groups this behavior separately from OWASP A09 security logging and alerting failures testing because the failure mode is different: logging gaps hide an attack after the fact, exception mishandling creates the attack surface itself. A verbose 500 error that reveals a stack trace with internal file paths, database schema names, or a third-party API key gives an attacker a map. A payment API that fails open when a fraud-check service times out gives an attacker a transaction.

For regulated industries, the business consequence is direct: an auditor asking for evidence of forced-failure testing, a board asking why a payment reconciliation broke during a vendor outage, or a customer asking why their session data appeared in another account after a server crash.

How to Test OWASP A10:2025 Mishandling of Exceptional Conditions

Testing exception handling requires deliberately breaking the application, not scanning it and waiting for a signature match. Follow this sequence for each in-scope service:

  1. Map the exception surface. Catalog every point where the application depends on external input, output, or a downstream service: API parameters, file uploads, third-party integrations, database calls, message queues, and webhook receivers.
  2. Force malformed and boundary input. Submit oversized payloads, null bytes, unexpected content types, negative numbers where positive integers are expected, and deeply nested JSON to trigger unhandled exceptions in parsing and validation layers.
  3. Simulate dependency failure. Kill or throttle a downstream dependency — a payment processor, an auth provider, a fraud-scoring API — mid-transaction and observe whether the application fails closed or fails open.
  4. Trigger resource exhaustion. Send concurrent requests that exhaust connection pools, memory, or file handles, then check whether the resulting error state exposes internal information or degrades access controls.
  5. Capture the full error response. Record HTTP status code, response body, headers, and timing for every forced error. A 500 response with a stack trace is a materially different finding than a generic 500 with an empty body.
  6. Validate business logic under exception. Interrupt a multi-step transaction — checkout, loan approval, claims submission — at each step and confirm the system doesn't commit a partial or inconsistent state.
  7. Check for information leakage in retries. Some applications behave correctly on the first forced failure but leak stack traces or internal identifiers on the second or third retry attempt.
  8. Retest after remediation. Confirm the fix closes the specific exception path tested, not just the one payload used to find it — chained exception paths are common across related endpoints.

API and Backend Exception Handling

APIs are the highest-yield surface for this category because every endpoint has to handle malformed requests from clients it doesn't control. During a manual API penetration test, testers submit malformed JSON, incorrect content-length headers, unsupported content types, and unexpected HTTP verbs against every endpoint, then compare the error response against what the API contract promises. A 500 error that returns a full stack trace, an ORM query fragment, or an internal hostname is a finding on its own — it hands an attacker the internal architecture without a single successful authenticated request.

Business Logic and Multi-Step Transactions

Exception mishandling in business logic is harder to find because it requires interrupting a workflow at a specific step, not sending one bad payload and reading the response. Killing a session mid-checkout, mid-refund, or mid-KYC verification and observing what state the system commits to is a manual test, not a scanner one. Chaining an interrupted transaction with a business logic flaw often produces the same category of result as testing for insecure design failures, because both depend on state a developer assumed would never occur — a refund that processes twice, a claim that submits without required documents, an approval that commits before the fraud check returns.

Resource Exhaustion and Fail-Open Conditions

Connection pool exhaustion, thread starvation, and memory pressure all produce exception states that development teams rarely test for directly. The test question is simple: when the system runs out of a resource, does it deny the request or grant it? A rate limiter that fails open under load, an authorization check that times out and defaults to "allow," or a queue that drops messages silently under backpressure are all A10:2025 findings with direct, quantifiable business impact.

Cloud and Infrastructure Exception Paths

Managed cloud services introduce their own exception surface: a Lambda function that times out mid-execution, an IAM policy evaluation that errors and defaults to permissive, or an autoscaling group that fails to provision under load and routes traffic to an unpatched fallback instance. Testers validate these paths by throttling API rate limits, forcing IAM policy evaluation errors, and observing what the infrastructure does when its normal control path is unavailable — the answer should always be deny, never allow.

AI and LLM Pipeline Exception Paths

Generative and agentic systems add a new exception surface: what happens when a model call times out, returns a malformed completion, or exceeds a token limit mid-response? A safe implementation fails to a deterministic, validated default. A vulnerable one executes a partial, unvalidated instruction or leaks the raw model error — including system prompt fragments — back to the user. This is a fast-growing test category for SaaS platforms shipping AI features in 2026, and it requires the same forced-failure methodology as classic API testing, applied to model calls instead of database calls.

Mobile and Client-Side Exception Handling

Mobile apps introduce client-side exception paths that don't exist server-side: an app that crashes cleanly on a malformed push notification is a stability bug; an app that writes an API key or refresh token to a device log during a crash is a security finding. OWASP's mobile testing guidance treats unhandled exceptions in local storage and inter-process communication handlers as a distinct check, separate from server-side error handling, and both need coverage in a full assessment.

Why Exception Handling Failures Vary in Severity

Not every unhandled exception carries the same risk. Severity depends on:

  • What the error response discloses. A stack trace with database credentials is critical; a generic "something went wrong" message is low severity.
  • Whether the failure is fail-open or fail-closed. An authorization check that grants access on timeout is far more severe than one that denies access on timeout.
  • How many steps the exception affects. A single malformed request that crashes one endpoint is lower severity than an exception that corrupts a multi-step transaction's state.
  • Whether the exception path is reachable pre-authentication. Unauthenticated attackers reaching a verbose error page multiplies exposure compared to a flaw only reachable after login.
  • How the application logs the event. A finding paired with a logging and alerting gap means the exploit attempt itself goes undetected, compounding the risk.
  • Whether remediation is systemic or a single patch. A fix that only closes the tested payload path, not the underlying missing input validation, leaves the same class of bug reachable elsewhere in the codebase.

Common Findings and Business Impact

Verbose stack trace on 500 error

  • Typical Root Cause: Missing global exception handler
  • Business Impact: Internal architecture disclosure, faster attack chaining

Fail-open authorization on timeout

  • Typical Root Cause: Default-allow logic inside try/catch
  • Business Impact: Unauthorized access during dependency outages

Partial transaction commit on interruption

  • Typical Root Cause: No rollback on exception
  • Business Impact: Financial reconciliation errors, data integrity disputes

Resource exhaustion crash

  • Typical Root Cause: No rate limiting or connection pool caps
  • Business Impact: Denial of service, SLA breach

Model error leaking system prompt

  • Typical Root Cause: No output sanitization on LLM failure
  • Business Impact: Prompt and IP disclosure, downstream prompt injection risk

Compliance Mapping for Exception Handling Failures

PCI DSS 4.0

  • What It Requires: Error handling must not disclose cardholder data or system details
  • What Assessors Check: Forced-error test evidence, error message content review

SOC 2

  • What It Requires: Availability and processing integrity criteria cover graceful degradation
  • What Assessors Check: Evidence that failures don't bypass access or processing controls

ISO 27001 Annex A (technological controls)

  • What It Requires: Secure coding and system acceptance testing must cover error conditions
  • What Assessors Check: Test cases for exception paths in the risk treatment plan

NIST CSF 2.0

  • What It Requires: Detect and Respond functions expect exception events to be logged, not silently dropped
  • What Assessors Check: Correlation between forced failures and monitoring alerts

HIPAA

  • What It Requires: Technical safeguards require that error states don't expose PHI
  • What Assessors Check: Error message review, breach-risk assessment of disclosed data

Testing this category in 2026 is no longer optional for regulated industries — auditors increasingly ask for evidence of forced-failure test cases as a deliverable, not just a clean automated scan report.

Manual Testing vs Automated Scanning for Exception Handling

Automated DAST/SAST scanning

  • Strengths: Fast, covers known error-pattern signatures at scale
  • Limitations: Can't interrupt multi-step transactions or simulate dependency failure
  • Best For: Baseline coverage, CI/CD gating

Manual penetration testing

  • Strengths: Forces real dependency failures, tests business logic under interruption, validates fail-open/fail-closed behavior
  • Limitations: Slower, requires skilled testers
  • Best For: Compliance evidence, high-risk transaction flows

Hacker-led red team simulation

  • Strengths: Chains exception paths with other vulnerability classes to prove exploitability
  • Limitations: Broader scope, higher cost
  • Best For: Board-level risk validation, pre-audit assurance

Automated scanners find the pattern; manual testers prove the exploit. A scanner flags a 500 error. A manual tester interrupts a loan disbursement at the exact moment the fraud-check API times out and shows the transaction clears anyway — that's the difference between a scan finding and a reportable business risk.

Common Mistakes When Testing This Category

  • Testing only the happy-path error. Sending one malformed payload and confirming a generic error message isn't sufficient — testers need to chain retries, concurrent requests, and combined failure conditions.
  • Treating a 500 error as automatically low severity. The status code says nothing about what the body or headers disclose; each response needs manual review.
  • Skipping dependency-failure simulation because it's disruptive. Staging environments exist for exactly this kind of forced-failure test; skipping it means the finding surfaces in production instead.
  • Assuming a fix works because the original payload no longer triggers it. Remediation verification needs to test the underlying validation gap, not just the one test case that found it.
  • Ignoring client-side and mobile exception paths. Server-side hardening doesn't protect against a mobile client that logs sensitive data to disk during a crash.

When and How Often to Test

Exception handling testing belongs in every full-scope penetration test, not as a standalone engagement. Teams shipping frequent releases in 2026 should validate this category during continuous penetration testing cycles rather than waiting for an annual assessment, since new dependencies and new failure modes ship with every deploy. High-risk workflows — payment processing, claims adjudication, identity verification — warrant a dedicated round of forced-failure testing any time the underlying integration or vendor changes, independent of the regular test calendar.

How AppSecure Security Tests for Mishandling of Exceptional Conditions

AppSecure Security treats exception handling as a first-class test category inside every penetration testing and red teaming engagement, not an afterthought bolted onto a vulnerability scan. Testers force dependency failures against payment gateways, identity providers, and internal microservices, then validate whether the application's fallback behavior is fail-closed by design or fail-open by accident.

For fintech, healthcare, and SaaS platforms carrying PCI DSS, HIPAA, or SOC 2 obligations, this test category produces the evidence auditors ask for directly: forced-error test cases, captured responses, and a remediation retest confirming the fix holds under the same conditions that surfaced it. Hacker-led testing goes further than a checklist — testers chain an exception-handling flaw with a broken access control or an injection point to show what an attacker actually gets, not just what could theoretically go wrong.

Validate your exception handling

Get manual test evidence for A10:2025 findings before your next audit.

Talk to AppSecure

Exception Handling Testing Checklist

  • Map every input, output, and dependency boundary in scope
  • Force malformed, oversized, and boundary-value input against each endpoint
  • Simulate downstream dependency timeout and failure
  • Capture full error response: status code, body, headers, timing
  • Test multi-step transactions for partial-commit and rollback behavior
  • Confirm authorization and rate-limiting logic fail closed, not open
  • Verify error messages exclude stack traces, internal paths, and credentials
  • Retest each fix against the original forced-failure condition

FAQ

What is OWASP A10:2025 Mishandling of Exceptional Conditions?

It's the OWASP category covering how applications behave when errors, timeouts, or resource limits occur — including verbose error disclosure, fail-open logic, and unhandled exceptions that corrupt transaction state. Testing it means deliberately forcing these conditions and evaluating the response, not just scanning for known error signatures.

How is exception handling testing different from vulnerability scanning?

Vulnerability scanning matches known signatures against a running application; exception handling testing manually forces failure conditions like dependency timeouts and interrupted transactions. Scanners rarely simulate a payment processor going down mid-transaction, which is exactly the condition that exposes fail-open logic.

Is mishandling of exceptional conditions the same as a denial-of-service vulnerability?

They overlap but aren't identical. A DoS vulnerability is about availability; mishandling of exceptional conditions covers availability, information disclosure, and authorization bypass triggered by the same forced-failure conditions.

How often should exception handling be tested?

Every full-scope penetration test should include it, and high-risk transaction flows should get a dedicated retest whenever an underlying vendor or integration changes. Annual testing alone misses failure modes introduced by frequent production releases.

Does PCI DSS require testing for error handling failures?

Yes. PCI DSS 4.0 requires that error handling not disclose cardholder data or internal system details, and assessors expect evidence from forced-error test cases as part of the penetration test scope.

Can automated tools find fail-open authorization bugs?

Rarely. Fail-open authorization bugs require simulating a dependency timeout during an active session, a multi-step condition automated scanners aren't built to reproduce reliably. Manual testing is the practical way to surface this class of finding.

What's the difference between A09 logging failures and A10 exception handling?

A09 covers whether security events get logged and alerted on after they happen. A10 covers whether the exception itself creates a vulnerability — a disclosure, a bypass, or a corrupted state — independent of whether it gets logged.

How long does testing this category take during a penetration test?

It's integrated throughout the engagement rather than timed separately, since exception paths get tested alongside API, business logic, and infrastructure testing. Scope, number of integrations, and transaction complexity determine the overall engagement timeline.

One Last Thing

The most damaging A10:2025 findings rarely show up in the first forced error — they show up in the third or fourth chained failure, when a system that fails closed once starts failing open under sustained load. Test exception handling under repeated, not single, failure conditions in 2026; one clean forced-error test proves nothing about what happens during a real outage or a sustained attack.

Related Guides

Tejas K. Dhokane, Marketing Associate at AppSecure Security
Tejas K. Dhokane

Tejas K. Dhokane is a marketing associate at AppSecure Security, driving initiatives across strategy, communication, and brand positioning. He works closely with security and engineering teams to translate technical depth into clear value propositions, build campaigns that resonate with CISOs and risk leaders, and strengthen AppSecure’s presence across digital channels. His work spans content, GTM, messaging architecture, and narrative development supporting AppSecure’s mission to bring disciplined, expert-led security testing to global enterprises.

Protect Your Business with Hacker-Focused Approach.

Loved & trusted by Security Conscious Companies across the world.
Stats

The Most Trusted Name In Security

450+
Companies Secured
7.5M $
Bounties Saved
4800+
Applications Secured
168K+
Bugs Identified
Accreditations We Have Earned
crest logo white
AICPA SOC 2 badge logo

Protect Your Business with Hacker-Focused Approach.