The Mathematics Search Engine

Mathematics News & Resources

4Mathematics is a specialist search engine for Mathematics. Discover the latest math news and mathematical content. Part of the 4SEARCH network of topic specific search engines.

Latest Articles


dev.to > renato_marinho > solving-the-non-deterministic-routing-problem-in-multi-tool-ai-agents-1l3n

Solving the Non-Deterministic Routing Problem in Multi-Tool AI Agents

14+ min ago   (464+ words) When building autonomous agents, we often treat tool calling as a black box. We provide a set of definitions to a Large Language Model (LLM) and assume that if the prompt is good enough, the model will pick the right function…...


marktechpost.com > 09/18/2026 > prismml-releases-ternary-bonsai-2-27b-a-5-9-gb-apache-2-0-model-retaining-98-2-of-qwen3-8-27b-performance

PrismML Releases Ternary Bonsai 2 27B: A 5.9 GB Apache 2.0 Model Retaining 98.2% of Qwen3.8 27B Performance

21+ min ago   (597+ words) Is it deployable? Yes. The Apache 2.0 weights run today on a 16 GB laptop or a single 24 GB GPU. You need PrismML’s llama.cpp fork or its MLX runtime. The model keeps the Qwen3.8 27B architecture unchanged. It has 27.36B parameters. That splits into…...


quantumzeitgeist.com > chapman-university-aumanns-theorem-gets-quantum

Aumann's Theorem Gets A Quantum Boost For All Generalized Probability Theories

7+ hour, 33+ min ago   (865+ words) Rusty is a quantum science nerd. He's been into academic science all his life, but spent his formative years doing less academic things. Now he turns his attention to write about his passion, the quantum realm. He loves all things…...


cryptobriefing.com > epoch-frontiermath-major-advance-gpt6-astra

Epoch classifies first AI solution as major advance in FrontierMath benchmark

46+ min ago   (311+ words) OpenAI's GPT-6 Astra scored nearly 98% on the hardest tier of math problems that stumped every AI model just 14 months ago When Epoch AI launched Tier 4 of its FrontierMath benchmark in July 2025, the best AI models on the planet could solve…...


dev.to > kruhlova > sound-analysis-has-no-cathiss-label-974

Sound Analysis Has No cat_hiss Label

1+ hour, 1+ min ago   (931+ words) Apple's built-in Sound Analysis model can label cat_meow and cat_purr. It has no cat_hiss class. Point a mic at a real hiss and you often get silence, a generic cat hit, or — worse — snake_hiss. Energy metering only proves something crossed a loudness line. Neither…...


note.com > sleepinglion0227 > n > n3bd064f719be

Mathematics makes AI lighter, and AI designs factories and science—5 latest AI papers|SleepingLion

2+ day, 2+ hour ago   (908+ words) Looking at today's papers, I feel that progress in AI research is no longer just about "making models even larger." This time, I have selected five papers focusing on mathematics, mathematics × AI, AI scientists, energy, industrial applications, and physical AI....


crn.com > news > software > 2026 > newly-merged-dbt-labs-and-fivetran-ready-channel-plans

Newly Merged Dbt Labs And Fivetran Ready Channel Plans For AI Data Push

1+ hour, 10+ min ago   (1260+ words) The newly created company, as yet unnamed, looks to be a force in the exploding AI data infrastructure arena with its combined data transformation and data integration technologies. The new company created through the recent merger of dbt Labs and…...


lesswrong.com > posts > xvdngZAqFZfek7KGH > pretraining-data-not-verifiability-is-why-llms-are

Pretraining data, not verifiability, is why LLMs are especially good at math (and coding) — LessWrong

1+ hour, 19+ min ago   (242+ words) Follow-up to: “LLMs are (still) mostly powered by imitative learning, not RL” A common take I’ve been hearing is: “LLMs are especially good at math[1] because math is easy to verify”. But that story doesn’t make much sense. So here’s…...


dzone.com > articles > when-your-benchmark-leaks-the-answer

When Your Benchmark Leaks the Answer

1+ hour, 27+ min ago   (894+ words) Bad synthetic data and flawed rules broke model evaluations, but high overall scores hid it. Test by category using realistic, production-style data. A detector I built was scoring 0.067 recall on temporal errors, meaning it caught about one in fifteen of…...