Analysis5 min read
The AI names the right approver. Did it use the right policy?
Why a €75,000 quote is a poor test of a changed approval threshold—and what a €40,000 case reveals.
In this article
A quote with a net value of €75,000 requires approval from the head of sales. That is the answer shown in our PULSE demonstration, and it agrees with the published policy. Yet this case tells us very little about whether the correct policy version was used.
The previous policy required approval above €25,000. The new threshold is €50,000. At €75,000, someone following the old rule would reach the same decision. A correct outcome can conceal a source-selection error.
What changed between the two versions
Our demonstration data shows version 3 of the Sales approval policy as published. The comparison displays the increase from €25,000 to €50,000 in net value and a reference to a decision dated 12 July 2026. Above the new threshold, the head of sales reviews the quote before it is sent. Up to and including €50,000, sales carries out the review independently.
PULSE demonstration data: the threshold changes from €25,000 to €50,000. This visible difference provides the basis for the comparison below.
This analysis considers that amount rule alone. It does not assume additional conditions such as special commercial terms. We have not run a language-model experiment here; the expected decisions can be derived directly from the two rules.
What four amounts can tell us
| Net value | Previous rule: above €25,000 | New rule: above €50,000 |
|---|---|---|
| €25,000 | Sales reviews independently | Sales reviews independently |
| €40,000 | Head-of-sales approval | Sales reviews independently |
| €50,000 | Head-of-sales approval | Sales reviews independently |
| €75,000 | Head-of-sales approval | Head-of-sales approval |
€40,000 falls between the thresholds. This case reveals whether the decision matches the old rule or the new one. Exactly €50,000 also separates the versions, while testing the meaning of “above”. Reading it as “at or above” produces the wrong approval route even when the number itself is correct.
The €75,000 case still has a purpose: it demonstrates the normal route to the head of sales. It simply answers a different question. A useful test set includes both ordinary cases and cases where competing explanations would produce different results.
A matching decision also needs matching evidence
Suppose the answer for a €40,000 quote correctly says that sales should review it independently. But the cited source is the previous policy with its €25,000 threshold. The decision is right; the evidence does not support it. Perhaps the new amount appeared in the question, or came from another source. The answer alone cannot establish what happened.
This separates the assessment of the decision from the assessment of its evidence. A reference is useful when the cited passage supports the particular claim. The number of documents linked says little about that.
The latest version may not be the applicable version
The comparison assumes that the new rule applies to the case. In an operational setting, that assumption needs to be established. A more recently saved draft may still await review. A published change may take effect next month. A historical quote may need to be assessed against the rule that applied at the time.
Version number, publication status and effective period answer different questions. A decision date also does not, by itself, establish retrospective applicability. We therefore do not infer a rule for historical transactions from the 12 July date shown in the demonstration.
A historical test needs an expectation confirmed by the responsible business expert: which date matters, and what transitional rule applies? If that has not been established, the assistance should identify this specific gap. Selecting the most recent document cannot substitute for the organisation’s decision.
One edited sentence can create several follow-up tasks
Publishing the knowledge page is not necessarily the last step. The process owner needs to establish whether an approval form or a rule in another application still uses €25,000. The person responsible for that system then implements any required change there. A link to the system shows a relationship; it does not confirm an update.
The same distinction matters when evaluating AI assistance. An answer saying “approval granted” and an approval status actually stored in the application are separate results to check. Anthropic’s evaluation methodology similarly distinguishes the conversation transcript from the resulting state of the environment.
Source [3]A small, maintained set of cases can cover these transitions: €40,000 for version selection, €50,000 for the boundary, an incomplete case for clarification and a case outside the established scope. Revisit the relevant cases after changes. A recorded test result should make clear whether it assessed only the answer or also an action that followed.
See how the policy, role and process fit together in the illustrated response. Walk through quote approval in PULSE
Sources & further reading
- PULSE: Knowledge and versioning
Source of the demonstration and the policy change analysed here. The comparison table is our own derivation from the two amount rules.
- Felix Drösel: Prozesse hier, Wissen dort
Background on applicability and responsibility for the use of knowledge; 2 September 2026. In German.
- Anthropic: Demystifying evals for AI agents
Methodological reference for distinguishing an answer from a verifiable outcome; 9 January 2026. It does not establish PULSE model performance.


