Skip to main content
Back to insights

INSIGHTS

Part 10 of 12

How to test missing exceptions and contradictory business rules

September 29, 2026•7 min read

A generated specification can be internally consistent and still miss the rule that changes an implementation. Testing the prose alone will not tell you whether the missing rule was available in a recording, never written down, or supplied only after someone asked the right person.

Product note: these essays explore the broader business-logic extraction category and future artifact possibilities. For the current RuleFoundry product, see the outputs page.

Illustration of business rules checked against evidence and unresolved questions

Core thesis

A useful evaluation separates those situations. It asks what information the system received, what it was allowed to ask, what it claimed, and whether an engineer could use the result without silently inventing business decisions.

Begin with a decision, not a document length

Consider this fictional purchasing policy: “Orders below $5,000 can be approved by the team manager.” A screen recording shows one manager approving a $2,400 order. A later policy note says restricted suppliers require procurement review.

Several decisions remain open. What happens at exactly $5,000? Does procurement review replace manager approval or come first? Who determines that a supplier is restricted? Does the new rule affect already approved orders?

The recording establishes an observed example. The policy introduces a condition. Neither settles every boundary or precedence question. A polished flowchart that silently chooses the answers is a failure even if its arrows look coherent.

For this constructed case, a useful initial result might say:

Manager approved the $2,400 order shown: Recording of that case; Observed behavior, not proof of every order's route.

Restricted suppliers require procurement review: Policy note; Supported written requirement, with applicability to confirm.

Exactly $5,000: Neither source states the boundary; Targeted follow-up.

Which approval takes precedence: Sources do not specify; Keep unresolved; do not invent an ordering.

That table is an expected reasoning pattern, not output we claim a product has produced. In an actual evaluation, preserve the original output and score it against a case definition fixed before the run.

Separate reading from eliciting

Use one lane to test whether a tool recovers decisions present in supplied material. Use another to test whether it obtains missing information from a standardized expert. Use a third to test a later correction that disagrees with the original source.

Do not score an interview tool as successful for recovering a fact you secretly gave it but withheld from another tool. Conversely, do not expect a file-only task to recover a fact that no source contains. All eligible tools should receive the same opportunity and budget to ask the expert in the elicitation lane.

The expert script should distinguish a known answer, an answer requiring another owner, and a genuinely unresolved decision. “I need to check with procurement” is valid evidence of a dependency. Turning it into a confident procurement rule is not a successful extraction.

Compare tools on the jobs they support

Requirements analysis is already a real capability. Kiro documents analysis of ambiguity, conflicting requirements and missing edge cases, followed by questions that can update the requirements. Generating specs and questions alone is not a useful claim of exclusivity. Kiro requirements analysis.

Workflow products also differ within the same vendor. Scribe Capture creates procedural guides, while Optimize documents issue analysis and conversational exploration. A fair evaluation names the module and task rather than assuming the entire product only records clicks. Scribe Capture, Optimize Agent.

Report an unsupported input path as a compatibility limitation. Do not disguise it as a zero on business-rule accuracy. Also measure a native workflow separately: setup, capture, correction and handoff effort can matter even when two products eventually reach a similar answer.

Score supported decisions and useful uncertainty

Keep the measurements separate:

Supported assertions: how many stated business rules are justified by admitted evidence or authorized expert answers? Show the numerator and denominator.

Recovery: how many relevant rules in the declared case key were captured under the allowed information and interview budget?

Unsupported assertions: which consequential conditions, defaults or precedence rules were invented? Show examples and their severity.

Useful follow-up: which questions changed understanding, established a missing dependency, or exposed a conflict? Extra questions without benefit cost time.

Traceability: does a citation support the claim, and does a correction preserve what changed? A valid timestamp alone does not prove the interpretation.

Downstream use: can another engineer implement and test the reviewed boundary while recognizing unresolved decisions?

A tool should not be rewarded for confidently completing every branch. Nor should it be rewarded for asking indefinitely. Measure expert time, operator edits and review effort alongside correctness. A concise explanation that already contains the necessary detail may need very little follow-up.

Keep the scorer independent of the proposed answer

For controlled fictional cases, write the case key before running the products. Keep it away from model-visible inputs. Make any hidden expert facts available only through the agreed interview mechanism, equally across eligible tools.

Have people review business correctness. Where practical, hide vendor identity during scoring and compare two reviewers' decisions on a sample. Record remaining disagreement instead of forcing a clean score. A model judge can assist with organization; it should not be the only authority for a vendor ranking.

For real organizational work, an answer key is often incomplete. Record who can approve a policy decision and which questions remain unresolved. Faithfulness to a source and correctness for the intended business are related, separate checks.

Preserve the baseline, failures and limits

Freeze product versions or run dates, inputs, prompts, expert access, time budgets and scoring criteria. Keep unsuccessful and interrupted runs. Use fresh sessions when behavior can vary, and report setup mistakes separately from product limits.

Improve the product on development cases, then evaluate the changed version on withheld cases. Do not tune against the withheld answers or present one attractive before-and-after example as general improvement. A small pilot can expose failure classes; it cannot establish broad market superiority.

Before using the result

  • Freeze the case, available information and expert-answer budget before comparison.
  • Keep supported rules, unsupported assertions and useful follow-up separate.
  • Retain failures and evaluate improvements on untouched cases.

Test the decisions an engineer would otherwise have to invent

RuleFoundry's hypothesis is that focused expert interviewing and review make it easier to obtain the business answers an implementation needs. This method is how that hypothesis should be challenged. The meaningful result is fewer unsupported decisions at a measured cost in expert and engineering effort—not simply a longer specification, a more fluent voice or a larger collection of documents.

This educational guide uses a constructed example, not a customer outcome or completed comparative benchmark. Product documentation was checked September 29, 2026.

Turn expert conversations into business logic your enterprise can trust

We will show you the rules, gaps, flows, and source trace that fall out once the logic is actually extracted and made reviewable.