# How to give AI agents exact math with a calculator tool

> LLMs guess at arithmetic, and floats round 0.1 + 0.2 wrong. Why agents miscalculate and how to give them an exact, safe calculator tool over MCP or REST.

Published 2026-10-04, updated 2026-10-04 by MCP Toolbelt. Canonical URL: https://mcptoolbelt.com/blog/calculator-tool-exact-math-ai-agents

## The short answer

An AI agent should not do arithmetic in its head. A large language model predicts the next token; it does not carry digits. It is usually right on small sums and confidently wrong on long multiplications, percentages of percentages and anything that needs exact rounding.

MCP Toolbelt's **`calculate`** tool does the arithmetic instead. It takes one argument, a math expression as a string, and returns the number:

- **Exact basic arithmetic.** `+`, `-`, `*`, `/`, `%` and integer powers run on exact rational numbers, so `0.1 + 0.2` is `0.3` and `1/3*3` is `1`.
- **Scientific functions.** `sqrt`, `cbrt`, `exp`, `ln`, `log`, `log2`, `log10`, `sin`, `cos`, `tan`, their inverses and hyperbolic forms, `atan2`, `hypot` and `pow`, plus the constants `pi`, `e` and `tau`.
- **Rounding you can rely on.** `round`, `floor`, `ceil` and `trunc` work on the exact value, and `round(x, 2)` rounds half away from zero.
- **Safe by design.** A small parser, not `eval`: no variables, no code execution and no network calls. Bad input returns a stable error code.

Agents call it over MCP (Model Context Protocol) or a plain REST endpoint, one expression per call or many in a batch.

## Why language models get math wrong

- **Models predict digits, they do not compute them.** `123456789 × 987654321` is `121932631112635269`. A model has to produce those 18 digits one token at a time, and a single slip in the middle looks just as plausible as the right answer.
- **Code is not automatically exact either.** An agent that writes `0.1 + 0.2` in Python or JavaScript gets `0.30000000000000004`, because binary floating point cannot store 0.1. Python's `round(2.675, 2)` returns `2.67` for the same reason.
- **Rounding rules differ.** Python rounds `-2.5` to `-2` (half to even); most invoices expect `-3`. Python says `-7 % 3` is `2`; C, Java, JavaScript and PHP say `-1`.
- **Operator precedence trips people and models up.** Is `-2^2` equal to `4` or `-4`? Is `3^2^2` equal to `81` or `6561`?
- **Answers are not reproducible.** Ask twice and you can get two totals. A tool returns the same number for the same expression every time.

Spinning up a code interpreter for every sum is slow, needs a sandbox and still inherits floating-point surprises. A calculator tool is one fast, deterministic call. The model decides _what_ to calculate; the tool does the calculation.

## What the calculate tool supports

| Feature        | Syntax                                                                                            | Exact?                         |
| -------------- | ------------------------------------------------------------------------------------------------- | ------------------------------ |
| Arithmetic     | `+` `-` `*` `/` and parentheses                                                                   | Yes                            |
| Remainder      | `%` (takes the sign of the dividend: `-7 % 3` is `-1`)                                            | Yes                            |
| Powers         | `^` or `**`, right-associative, binding tighter than unary minus: `-2^2` is `-4`, `3^2^2` is `81` | Yes for integer exponents      |
| Numbers        | `42`, `3.14`, `.5`, `1.5e3`                                                                       | Yes                            |
| Rounding       | `round(x)`, `round(x, digits)` with `digits` from −308 to 308, `floor`, `ceil`, `trunc`           | Yes                            |
| Other exact    | `abs`, `min(a, b, ...)`, `max(a, b, ...)`, `pow(x, y)` (the same as `x^y`)                        | Yes                            |
| Roots and logs | `sqrt`, `cbrt`, `exp`, `ln`, `log(x)` (base 10), `log(x, base)`, `log2`, `log10`, `hypot(x, y)`   | Approximate (double precision) |
| Trigonometry   | `sin`, `cos`, `tan`, `asin`, `acos`, `atan`, `atan2(y, x)`, `sinh`, `cosh`, `tanh`, in radians    | Approximate (double precision) |
| Constants      | `pi`, `e`, `tau`                                                                                  | Approximate (double precision) |

Function names are case-insensitive and whitespace is ignored. An expression can be up to 1,000 characters long.

The output is `{"expression": "...", "result": ...}`. `result` is a JSON number: an integer when the exact result is a whole number that fits in a signed 64-bit integer, otherwise the nearest double. Exact intermediate values may grow to 4,096 digits.

## Calculate an expression with one API call

What is `123456789 × 987654321`?

```bash
curl -s https://mcptoolbelt.com/v1/tools/calculate \
  -H "Authorization: Bearer $MCP_TOOLBELT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"expression": "123456789 * 987654321"}'
```

The result (the REST API wraps it in `data.output`):

```json
{
    "expression": "123456789 * 987654321",
    "result": 121932631112635269
}
```

Every digit is right, and it is an integer, not `1.2193263111263526e17`. Whole numbers up to 9,223,372,036,854,775,807 (the largest signed 64-bit integer) come back as integers. `2^63` and larger do not fit and come back as the double `9.223372036854776e+18`.

## Exact decimals: why 0.1 + 0.2 is 0.3 here

```json
{ "expression": "0.1 + 0.2" }
```

returns `0.3`. Basic arithmetic runs on exact fractions, so `0.1` really is one tenth and the sum really is three tenths. Only the final result is converted to a JSON number. Likewise:

| Expression        | `calculate` result | Floating point (Python, JavaScript) |
| ----------------- | ------------------ | ----------------------------------- |
| `0.1 + 0.2`       | `0.3`              | `0.30000000000000004`               |
| `0.07 * 100`      | `7`                | `7.000000000000001`                 |
| `1.1 * 3`         | `3.3`              | `3.3000000000000003`                |
| `round(2.675, 2)` | `2.68`             | `2.67`                              |
| `round(-2.5)`     | `-3`               | `-2` (Python)                       |

This matters for money. Rounding the wrong way on one line of an invoice is a one-cent difference; on a thousand lines it is a reconciliation problem.

## Everyday calculations for agents

### Prices, VAT and discounts

Three items at €49.99 with 21% VAT:

```json
{ "expression": "round(49.99 * 3 * 1.21, 2)" }
```

returns `181.46`. Without `round` the result is `181.4637`. Put the rounding in the expression, where your business rule says it belongs (per line or per total), instead of asking the model to round afterwards.

Note that `%` is the remainder operator, not "percent". Write 12.5% of 80 as `80 * 12.5 / 100` or `80 * 0.125` (both `10`); `12.5% * 80` is rejected as `invalid_expression`.

### Compound interest

€1,000 at 5% a year, compounded monthly, for 10 years:

```json
{ "expression": "round(1000 * (1 + 0.05/12)^(12*10), 2)" }
```

returns `1647.01`. The power has an integer exponent, so it is computed exactly before rounding; the unrounded result is `1647.009497690283`.

### Unit-free engineering math

`hypot(3, 4)` is `5`, `log(8, 2)` is `3`, `(2 + 3) * sqrt(16) / 4` is `5` and `atan2(1, 1) * 4` is `3.141592653589793`. To convert between units such as miles and kilometres, use the separate `convert_units` tool rather than typing conversion factors into an expression.

## Where results are approximate

Square roots, logarithms, trigonometry and fractional powers such as `2^0.5` use IEEE 754 double precision, which is about 15 to 17 significant digits. The tool never pretends otherwise:

- `sin(pi)` returns `1.2246467991473532e-16`, not `0`, because `pi` itself is a 17-digit approximation. Round the result if you need a tidy number.
- Exact arithmetic resumes after the function: `sqrt(16)` is `4`, and everything around a function call keeps its precision.
- Results that are too large for a double (`10^400`) fail with `overflow`, and nonzero results too small to represent (`exp(-1000)`) fail with `underflow`, instead of silently becoming infinity or zero.

## Use it from an MCP client

Any client that supports remote MCP servers (Claude, Claude Code, Cursor, VS Code with GitHub Copilot, ChatGPT developer mode, Gemini CLI, Codex and others) can connect to MCP Toolbelt with one URL:

```text
https://mcptoolbelt.com/mcp
```

Once connected, `calculate` shows up next to the other tools with its input and output schema, and the model calls it with the same `expression` argument as the REST API. Tell the agent when to use it:

```text
Never do arithmetic yourself, not even simple sums. Call the calculate tool
with one expression and use its result. Put any rounding rule in the
expression, e.g. round(subtotal * 1.21, 2). Write percentages as
multiplication: 15% of x is x * 0.15. If it returns an error, fix the
expression instead of estimating.
```

Agents without an account can register themselves through [/auth.md](https://mcptoolbelt.com/auth.md). Prices for every tool are published at [/pricing.json](https://mcptoolbelt.com/pricing.json) and in the [tool catalog](https://mcptoolbelt.com/tools).

## Run many calculations in one batch

Totalling every line of an order or recomputing a column of a spreadsheet? When batching is enabled for the tool, send many expressions to `/v1/tools/calculate/batch`:

```json
{
    "inputs": [
        { "expression": "round(49.99 * 3 * 1.21, 2)" },
        { "expression": "round(19.99 * 0.85, 2)" },
        { "expression": "round(1000 * (1 + 0.05/12)^(12*10), 2)" }
    ]
}
```

Results come back in the same order in `data.output.results`: `181.46`, `16.99` and `1647.01`. A calculation either produces a number or fails, so one failing item (a division by zero, say) fails the whole batch, the error names the item, and nothing is charged. Send an `Idempotency-Key` header so retries are never charged twice.

## Error codes your agent can branch on

| Code                 | Meaning                                                                        |
| -------------------- | ------------------------------------------------------------------------------ |
| `invalid_input`      | The arguments do not match the schema, such as a missing or empty `expression` |
| `invalid_expression` | A syntax error, with its position: `2 +` gives "Unexpected end of expression." |
| `unknown_identifier` | An unknown function or constant, such as `foo(2)` or a variable name `x`       |
| `division_by_zero`   | Division or remainder by zero, including `0^-1`                                |
| `domain_error`       | Undefined for these values, such as `sqrt(-1)` or `ln(0)`                      |
| `overflow`           | The result is too large to represent, such as `10^400`                         |
| `underflow`          | A nonzero result too small to represent, such as `exp(-1000)`                  |
| `calculation_limit`  | Exact arithmetic would exceed 4,096 digits, such as `10^5000`                  |

Calls that fail with one of these errors are not charged.

## Frequently asked questions

### Can ChatGPT or Claude do math on their own?

Not reliably. Language models generate numbers as text and often get long multiplications, divisions and rounding wrong while sounding certain. Connect a calculator tool such as `calculate` over MCP and let the model call it for every calculation.

### Why does 0.1 + 0.2 equal 0.30000000000000004 in code?

Most programming languages store decimals as binary floating-point numbers, which cannot represent 0.1 or 0.2 exactly. The `calculate` tool does basic arithmetic on exact fractions, so `0.1 + 0.2` returns `0.3`.

### Is it safe to send user input to the calculator?

Yes. The expression is parsed by a small dedicated grammar, never run with `eval` or a code interpreter. It accepts only numbers, operators, parentheses and a fixed list of functions and constants, and it limits length, nesting depth and the size of exact numbers.

### Does it support variables or equations?

No. An expression contains only numbers, constants and functions, so the agent substitutes the values itself: `round(1000 * (1 + 0.05/12)^120, 2)`, not `P * (1 + r/n)^(n*t)`. It does not solve equations or do symbolic algebra.

### How does round work, and are angles in degrees?

`round(x, digits)` rounds half away from zero on the exact value, so `round(2.675, 2)` is `2.68` and `round(-2.5)` is `-3`. Negative `digits` round to tens, hundreds and so on: `round(1234567, -3)` is `1235000`. Trigonometric functions use radians; convert degrees with `round(sin(30 * pi / 180), 12)`, which is `0.5` (without `round` it is `0.49999999999999994`).

### What does % mean in an expression?

The remainder after division, with the sign of the dividend: `7 % 3` is `1` and `-7 % 3` is `-1`. It is not a percent sign; write 15% of x as `x * 0.15`.

### What does a calculation cost?

Each expression is one call unit. Calls that fail because of invalid input are not charged. Current prices and any free monthly calls are listed in the [tool catalog](https://mcptoolbelt.com/tools) and at [/pricing.json](https://mcptoolbelt.com/pricing.json).

## Sources

- [IEEE 754](https://standards.ieee.org/ieee/754/6210/): the floating-point standard behind doubles and the 0.1 + 0.2 problem.
- [Python documentation, Floating-Point Arithmetic: Issues and Limitations](https://docs.python.org/3/tutorial/floatingpoint.html): why `round(2.675, 2)` gives `2.67`.
- [What Every Computer Scientist Should Know About Floating-Point Arithmetic](https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.html) by David Goldberg.
- [Model Context Protocol](https://modelcontextprotocol.io/): the open standard agents use to discover and call tools.
