Markdown for LLMs
Science tasks: models with and without tools
The source Markdown for this article. Copy it into your assistant or download it as a text file.
# Science tasks: models with and without tools
**Result:** on 1 October 2026, three models solved three science tasks
without tools. The claim that a plain model fails without tools was **not
confirmed** on these tasks.
- `openai/gpt-6-luna` (Luna) solved the elliptic-curve and homology tasks.
On the uncertainty budget it computed correctly and noted the rounding of
a table coefficient.
- `qwen/qwen3-30b-a3b-instruct-2507` reached correct final answers on the
first two tasks after correcting its own intermediate errors.
- `google/gemini-2.5-flash-lite` made errors.
For Luna on these sizes, the Arxo tools were not needed to reach the right
mathematical answer. What they add is a different product, described
below. There was one run per task and mode, so this is not a ranking.
## Tasks
- **Elliptic-curve chain.** On y² = x³ + 2x + 3 over GF(97), check curve
membership of four points, two supplied modular inverses and two slopes,
and two proposed additions. Two separate variants use a wrong sum and a
wrong inverse.
- **Homology certificate.** For two integer boundary maps, check that their
product is zero in all 16 entries, that a proposed cycle is a cycle and
that a proposed filling maps to it. Then check a corrupted filling
coordinate by coordinate, and compute a minor and the rank lower bound it
gives.
- **Uncertainty budget (GUM).** For a sum of two independent inputs with
standard uncertainties of 3 µV and 4 µV and degrees of freedom 4 and 6,
compute:
- the combined variance and the combined standard uncertainty;
- the Welch-Satterthwaite effective degrees of freedom and their
truncation;
- the 95 % coverage factor from table G.2 and the expanded uncertainty.
## Setup
**Without tools.** One direct request per model and task through
OpenRouter, with no tools field. The limit was 12,000 output tokens and 300
seconds. Luna and Gemini used medium reasoning. The Qwen instruct model
received no reasoning setting, since it does not support one. The original
numbers and proposed certificates were kept. Requirements to call the
tools, to use only engine answers and to report the engine's "not
established" status were removed. The model could compute on its own but
was not allowed to claim tool execution. No engine answers, expected
numbers or solution hints were given to the models.
**With tools.** For Luna, earlier successful runs of the same tasks with
the real Arxo tool server were reused as controls. They used 24 tools, with
a deadline of 180 seconds for homology and the uncertainty budget and 300
seconds for the elliptic chain. All finished within the limit. For Qwen and
Gemini, new tool attempts were made for homology and the uncertainty budget.
The elliptic task with tools had been run earlier the same day.
**Checking.** Final answers and uncorrected intermediate steps were checked
by hand against the saved tool results and independent integer and rational
arithmetic. An explicit correction of one's own error was allowed. "Solved
the mathematical task correctly" was scored separately from "obtained an
executable engine certificate with an audit of the premises".
## Results without tools
| Model | Elliptic chain | Homology | Uncertainty budget | Cost of three answers |
|---|---|---|---|---:|
| Luna | Correct | Correct | Correct, with a note on table rounding | $0.0032571 |
| Qwen3 30B A3B Instruct 2507 | Correct after self-correction | Correct after self-correction | Main calculation correct, but does not match the given table | $0.0031667 |
| Gemini 2.5 Flash Lite | Wrongly rejects P + P = Q | Main answers correct, one uncorrected error in an intermediate result | Wrong formula for degrees of freedom | $0.0083474 |
Details checked:
- **Elliptic chain.** Expected: P, Q and R on the curve, S not; inverses 89
and 63; slopes 59 and 58; P + P = (80, 10); P + Q = (80, 87); the
inverse 62 is wrong.
- Luna and Qwen got all of this. Qwen first multiplied 97 × 66 wrongly,
then corrected the remainder and the final answer.
- Gemini wrote 59² = 35 × 97 + 76, where the correct remainder is 86. It
got x(2P) = 70 instead of 80 and rejected the correct doubling.
- **Homology.** Expected: the product of the boundary maps is zero in all
16 entries; the cycle and the filling check out; the minor is 1; the rank
lower bound is 2; the cycle is a boundary. With the corrupted filling the
image is (−498, −210, 210, 72), and no coordinate matches.
- Luna got all of this.
- Qwen first wrote −526, then corrected it to −498.
- Gemini kept −598 instead of −498, although its "no match" conclusion
was right. It also described being a cycle as enough for
null-homology.
- **Uncertainty budget.** Expected: variances 25 and 37; standard
uncertainties 5 and 6.08; effective degrees of freedom 1500/151 ≈
9.9337748, truncated to 9. The table G.2 used in the Arxo package gives
k = 2.26 and U = 11.30.
- Luna first used the more precise quantile 2.262 and got 11.31. It then
gave the table values 2.26 and 11.30 separately and explained the
difference.
- Qwen used 2.262 and 11.31 without that note and called it the exact
table value. Its approximate 9.922 for the degrees of freedom was also
inexact, although the truncation was right. This is a mismatch with the
chosen table, not an error in the propagation law.
- Gemini divided the contributions by νᵢ − 1 instead of νᵢ. It got a
truncation of 7, k = 2.365 and U = 11.83, which is a method error.
## Luna: observed cost and time
| Task | With tools, saved successful run | Without tools, new run |
|---|---:|---:|
| Elliptic chain | $0.0032216 · 38.7 s | $0.0010782 · 17.4 s |
| Homology | $0.0134214 · 61.0 s | $0.0008822 · 12.6 s |
| Uncertainty budget | $0.0046258 · 59.0 s | $0.0012967 · 22.0 s |
| Total | $0.0212689 · 158.8 s | $0.0032571 · 52.0 s |
In these observations Luna without tools was about 6.5 times cheaper and
about 3 times faster over the summed sequential times. This is not a fixed
rate and not a clean ablation of one flag. The tool runs were made earlier,
with a larger system context, tool catalogues and several rounds. Routing
and cache state were not aligned. Costs are the reported usage, including
billed reasoning tokens, not list prices.
## What the tools add
For these sizes and numbers, the tools did not change whether Luna reached
the right mathematical answer. They produce a different result:
- an executable, independent check of each claim;
- the provenance of every number;
- a reproducible set of facts;
- an explicit limit of what is supported;
- a distinction between "not confirmed" and "refuted".
## Tool runs of the other two models
In this run, tool use by Qwen and Gemini in the shared conversational setup
was unstable. In the earlier elliptic run, Qwen did not pass facts and
Gemini invented predicates. The new attempts also had provider refusals,
exhausted output tokens and answers without a calculation call. These
failures are recorded separately. They do not show that the model cannot
do the mathematics: Qwen solved the first two tasks without tools.
- **Qwen, homology, repeat.** The engine confirmed the zero product, the
cycle, the original filling and null-homology. Wrong initial calls used
up the time: 300 seconds, 20 calls, $0.0177066. The corrupted variant was
checked in two coordinates only, the minor was not computed, and there
was no final answer. This is a partial tool result.
- **Qwen, uncertainty budget, repeat.** Reached a combined variance of 25,
5 µV, 9 degrees of freedom and exactly the table factor 2.26, after type
errors and empty selections. It took 24 calls, 165.1 seconds and
$0.0294043. The expanded uncertainty and the correlated variant were not
established, and there was no final answer. Here the tool fixed the table
mismatch, but the path was costly and incomplete.
- **Gemini, homology.** Read two catalogues, made no query call and ran out
of output tokens: 78.9 s, $0.0082734.
- **Gemini, uncertainty budget, repeat.** Ended with a short message about
looking for predicates, without a tool call: 21.4 s, $0.003302.
No successful tool solution was obtained from Gemini in this run. The
internal cause of the initial API stream failures was not established.
## Limits
- One run per task and mode, with separate retries only for API failures.
Not a statistical ranking of models.
- Three tasks of small size. Larger matrices, long traces and random
parameter sets were not measured. These three examples do not establish a
limit of what a plain model can do.
- The Luna tool controls were recorded earlier and under a different
context, so the cost and time comparison is not synchronous.
- The session made 16 new requests: 9 without tools, 4 with tools and 3
repeats. Their reported usage sums to $0.0762894. This is not a billing
statement: an interrupted round may be missing from the usage.
- The main application model and its flow were not changed for the run.