The Useful Question
English

How to check an AI-built calculator with one awkward discount

A calculator returns 960 for a price of 1,200 with 20% off. That looks right. Now try a price of 101 with 50% off. Depending on what gets rounded first, two plausible implementations return different answers.

Write down the expected answer before checking the code. Then choose an example that forces the rounding rule to matter. This tutorial uses a deliberately faulty calculation and a corrected version, with runnable tests. It is a teaching example, not a benchmark of a particular AI coding product.

Decide what the calculator promises

For this example, the contract is deliberately narrow:

A minor unit is the unit you choose to count: cents, for example. A price of 101 cents means 1.01 in its larger currency unit. The calculation works with the integer 101. It does not assume every currency has the same subdivision.

The rounding rule is our specification, not a universal retail rule. If your intended use requires rounding the discount instead, define that different contract first.

Microsoft recommends reading and testing generated code, including checking input types and ranges. This example turns that advice into observable results. Microsoft Learn

Make the answer independent of the implementation

On paper, 101 × 50 ÷ 100 is 50.5. Our rule makes the payable amount 51 minor units. Save that expected result before running the function. Copying the function's output into the test's answer field would merely preserve its mistake.

Use these nearby cases to expose where rounding changes the result:

Scroll the table horizontally; arrow keys work when focused.
Price in minor unitsDiscountUnrounded payable amountExpected result
1,20020%960960
10150%50.551
151%0.490
150%0.501
149%0.511

The one-unit prices are useful because they straddle the boundary. They do not need to resemble typical purchases.

Introduce one convincing mistake

The faulty version calculates and rounds the discount first:

Discount: 101 × 50 ÷ 100 = 50.5 → 51

Payable amount: 101 − 51 = 50

Its subtraction is correct, but it violates the chosen contract. The ordinary 1,200-at-20% case returns 960 in both versions, so that case alone cannot reveal the mistake.

The corrected calculation rounds the payable amount directly:

Payable amount = integer part of (price × (100 − discount percentage) + 50) ÷ 100

For 101 at 50%, this gives (5,050 + 50) ÷ 100 = 51. The source uses BigInt integer arithmetic; its division discards the fractional part. With our nonnegative inputs, adding 50 before division implements the chosen half-up rule. Numeric inputs are checked with Number.isSafeInteger before conversion. MDN BigInt, MDN Number.isSafeInteger

Run the same tests against both versions

On October 10, 2026, the calculation functions were run locally with Node.js v24.19.0. The suite contains 35 cases covering ordinary values, rounding boundaries, the amount cap, 0% and 100%, invalid arguments, and repeated calculation.

Two failures concern the 101-unit and one-unit half-price cases. The third repeats the same rounding defect within a sequence; it is not a separate newly discovered bug.

That sequence calculates 1,200 at 20%, then 101 at 50%, then the original input again. Expected results stay 960 → 51 → 960. Another case checks that a valid calculation still works after rejecting a blank string.

Keep the expected results unchanged while fixing the implementation. GitHub's review guidance specifically warns about failing tests being deleted or skipped. GitHub Docs

The calculation source, test source, actual execution log and reproduction instructions are included.

What passing establishes

These tests cover functions, not a working app. No browser input, text-to-number conversion, error display, repeated button clicks or mobile layout was tested. Passing 35 cases neither proves every input correct nor establishes readiness to publish.

For your app, the next step is to enter the same examples through its actual interface and compare the displayed results. Preserve three things: the expected answers, evidence of what ran, and what remains unchecked. That gives you a concrete basis for deciding what to fix or test next.