# Science tasks: models with and without tools **Result:** on 1 October 2026, three models solved three science tasks without tools. The claim that a plain model fails without tools was **not confirmed** on these tasks. - `openai/gpt-6-luna` (Luna) solved the elliptic-curve and homology tasks. On the uncertainty budget it computed correctly and noted the rounding of a table coefficient. - `qwen/qwen3-30b-a3b-instruct-2507` reached correct final answers on the first two tasks after correcting its own intermediate errors. - `google/gemini-2.5-flash-lite` made errors. For Luna on these sizes, the Arxo tools were not needed to reach the right mathematical answer. What they add is a different product, described below. There was one run per task and mode, so this is not a ranking. ## Tasks - **Elliptic-curve chain.** On y² = x³ + 2x + 3 over GF(97), check curve membership of four points, two supplied modular inverses and two slopes, and two proposed additions. Two separate variants use a wrong sum and a wrong inverse. - **Homology certificate.** For two integer boundary maps, check that their product is zero in all 16 entries, that a proposed cycle is a cycle and that a proposed filling maps to it. Then check a corrupted filling coordinate by coordinate, and compute a minor and the rank lower bound it gives. - **Uncertainty budget (GUM).** For a sum of two independent inputs with standard uncertainties of 3 µV and 4 µV and degrees of freedom 4 and 6, compute: - the combined variance and the combined standard uncertainty; - the Welch-Satterthwaite effective degrees of freedom and their truncation; - the 95 % coverage factor from table G.2 and the expanded uncertainty. ## Setup **Without tools.** One direct request per model and task through OpenRouter, with no tools field. The limit was 12,000 output tokens and 300 seconds. Luna and Gemini used medium reasoning. The Qwen instruct model received no reasoning setting, since it does not support one. The original numbers and proposed certificates were kept. Requirements to call the tools, to use only engine answers and to report the engine's "not established" status were removed. The model could compute on its own but was not allowed to claim tool execution. No engine answers, expected numbers or solution hints were given to the models. **With tools.** For Luna, earlier successful runs of the same tasks with the real Arxo tool server were reused as controls. They used 24 tools, with a deadline of 180 seconds for homology and the uncertainty budget and 300 seconds for the elliptic chain. All finished within the limit. For Qwen and Gemini, new tool attempts were made for homology and the uncertainty budget. The elliptic task with tools had been run earlier the same day. **Checking.** Final answers and uncorrected intermediate steps were checked by hand against the saved tool results and independent integer and rational arithmetic. An explicit correction of one's own error was allowed. "Solved the mathematical task correctly" was scored separately from "obtained an executable engine certificate with an audit of the premises". ## Results without tools | Model | Elliptic chain | Homology | Uncertainty budget | Cost of three answers | |---|---|---|---|---:| | Luna | Correct | Correct | Correct, with a note on table rounding | $0.0032571 | | Qwen3 30B A3B Instruct 2507 | Correct after self-correction | Correct after self-correction | Main calculation correct, but does not match the given table | $0.0031667 | | Gemini 2.5 Flash Lite | Wrongly rejects P + P = Q | Main answers correct, one uncorrected error in an intermediate result | Wrong formula for degrees of freedom | $0.0083474 | Details checked: - **Elliptic chain.** Expected: P, Q and R on the curve, S not; inverses 89 and 63; slopes 59 and 58; P + P = (80, 10); P + Q = (80, 87); the inverse 62 is wrong. - Luna and Qwen got all of this. Qwen first multiplied 97 × 66 wrongly, then corrected the remainder and the final answer. - Gemini wrote 59² = 35 × 97 + 76, where the correct remainder is 86. It got x(2P) = 70 instead of 80 and rejected the correct doubling. - **Homology.** Expected: the product of the boundary maps is zero in all 16 entries; the cycle and the filling check out; the minor is 1; the rank lower bound is 2; the cycle is a boundary. With the corrupted filling the image is (−498, −210, 210, 72), and no coordinate matches. - Luna got all of this. - Qwen first wrote −526, then corrected it to −498. - Gemini kept −598 instead of −498, although its "no match" conclusion was right. It also described being a cycle as enough for null-homology. - **Uncertainty budget.** Expected: variances 25 and 37; standard uncertainties 5 and 6.08; effective degrees of freedom 1500/151 ≈ 9.9337748, truncated to 9. The table G.2 used in the Arxo package gives k = 2.26 and U = 11.30. - Luna first used the more precise quantile 2.262 and got 11.31. It then gave the table values 2.26 and 11.30 separately and explained the difference. - Qwen used 2.262 and 11.31 without that note and called it the exact table value. Its approximate 9.922 for the degrees of freedom was also inexact, although the truncation was right. This is a mismatch with the chosen table, not an error in the propagation law. - Gemini divided the contributions by νᵢ − 1 instead of νᵢ. It got a truncation of 7, k = 2.365 and U = 11.83, which is a method error. ## Luna: observed cost and time | Task | With tools, saved successful run | Without tools, new run | |---|---:|---:| | Elliptic chain | $0.0032216 · 38.7 s | $0.0010782 · 17.4 s | | Homology | $0.0134214 · 61.0 s | $0.0008822 · 12.6 s | | Uncertainty budget | $0.0046258 · 59.0 s | $0.0012967 · 22.0 s | | Total | $0.0212689 · 158.8 s | $0.0032571 · 52.0 s | In these observations Luna without tools was about 6.5 times cheaper and about 3 times faster over the summed sequential times. This is not a fixed rate and not a clean ablation of one flag. The tool runs were made earlier, with a larger system context, tool catalogues and several rounds. Routing and cache state were not aligned. Costs are the reported usage, including billed reasoning tokens, not list prices. ## What the tools add For these sizes and numbers, the tools did not change whether Luna reached the right mathematical answer. They produce a different result: - an executable, independent check of each claim; - the provenance of every number; - a reproducible set of facts; - an explicit limit of what is supported; - a distinction between "not confirmed" and "refuted". ## Tool runs of the other two models In this run, tool use by Qwen and Gemini in the shared conversational setup was unstable. In the earlier elliptic run, Qwen did not pass facts and Gemini invented predicates. The new attempts also had provider refusals, exhausted output tokens and answers without a calculation call. These failures are recorded separately. They do not show that the model cannot do the mathematics: Qwen solved the first two tasks without tools. - **Qwen, homology, repeat.** The engine confirmed the zero product, the cycle, the original filling and null-homology. Wrong initial calls used up the time: 300 seconds, 20 calls, $0.0177066. The corrupted variant was checked in two coordinates only, the minor was not computed, and there was no final answer. This is a partial tool result. - **Qwen, uncertainty budget, repeat.** Reached a combined variance of 25, 5 µV, 9 degrees of freedom and exactly the table factor 2.26, after type errors and empty selections. It took 24 calls, 165.1 seconds and $0.0294043. The expanded uncertainty and the correlated variant were not established, and there was no final answer. Here the tool fixed the table mismatch, but the path was costly and incomplete. - **Gemini, homology.** Read two catalogues, made no query call and ran out of output tokens: 78.9 s, $0.0082734. - **Gemini, uncertainty budget, repeat.** Ended with a short message about looking for predicates, without a tool call: 21.4 s, $0.003302. No successful tool solution was obtained from Gemini in this run. The internal cause of the initial API stream failures was not established. ## Limits - One run per task and mode, with separate retries only for API failures. Not a statistical ranking of models. - Three tasks of small size. Larger matrices, long traces and random parameter sets were not measured. These three examples do not establish a limit of what a plain model can do. - The Luna tool controls were recorded earlier and under a different context, so the cost and time comparison is not synchronous. - The session made 16 new requests: 9 without tools, 4 with tools and 3 repeats. Their reported usage sums to $0.0762894. This is not a billing statement: an interrupted round may be missing from the usage. - The main application model and its flow were not changed for the run.