The Mathematics Search Engine

Mathematics News & Resources

4Mathematics is a specialist search engine for Mathematics. Discover the latest math news and mathematical content. Part of the 4SEARCH network of topic specific search engines.

Latest Articles


lesswrong.com > posts > HJPJArDvcvRBkbaEa > narrow-multimodal-fine-tuning-can-induce-emergent

Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment — LessWrong

26+ min ago   (821+ words) As multimodal systems are used more and more widely, their emergent misalignment (EM) behavior (Betley et al., 2025) also deserves attention. We used three narrow training tasks: For open models we train rank-32 LoRA adapters (Hu et al., 2022) on the language-model…...


dev.to > ayush_pangaonkar > 965-accurate-spam-filter-and-the-58-spam-messages-it-let-through-akm

96.5% Accurate Spam Filter, and the 58 Spam Messages It Let Through

1+ day, 8+ min ago   (367+ words) My spam classifier scores 96.5% accuracy on the test set. That sounds great. The confusion matrix tells a more useful story. The dataset is mail_data.csv, with 5,572 messages: 4,825 ham (86.6%) and 747 spam (13.4%), no missing values. I mapped labels to 0 (spam) and 1 (ham), split…...


dev.to > ayush_pangaonkar > my-near-perfect-r2-was-fake-debugging-a-speed-and-distance-model-43mn

My Near-Perfect R2 Was Fake: Debugging a Speed and Distance Model

16+ hour, 57+ min ago   (335+ words) The result looked great: R2 of 0.9993 and a correlation of 0.9994 between speed and distance. The problem was that the data had no real relationship in it. Speed and distance were generated independently, so by construction there should be no relationship between…...


dev.to > carbonlayer > a-per-inference-number-is-not-a-healthcare-emissions-inventory-2olf

A Per-Inference Number Is Not a Healthcare Emissions Inventory

12+ min ago   (1062+ words) Production AI teams already track latency, token usage, error rates, and cost. The next question is increasingly environmental: How much energy did this inference use? What carbon impact can be attributed to it? Can water use be estimated? Which values…...


dev.to > carbonlayer > what-if-every-ai-inference-came-with-a-transparent-impact-receipt-15b1

What If Every AI Inference Came With a Transparent Impact Receipt?

12+ min ago   (392+ words) An AI API gives you the model’s answer. Often, it also gives you token counts. But when you’re running inference in production, you may also need to know: What did that request cost? How long did it take? And what…...


dev.to > garje > google-hotels-api-hotel-prices-by-date-as-json-123

Google Hotels API: hotel prices by date as JSON

14+ min ago   (706+ words) Google Hotels shows prices from many booking sites for any city and night. If you want those prices as data, for a price calendar, a daily tracker or a rate check, here is what I learned collecting them, and the…...


dev.to > selfhostpilot > how-much-ram-you-actually-need-to-run-ai-locally-5hl2

How much RAM you actually need to run AI locally

20+ min ago   (278+ words) Higher-end AI PCs landed this month at prices that make people ask the wrong question first: which one is worth the money? The number that actually decides what you can run is not the price, the core count, or the…...


dev.to > hawkbtcommander > my-backtest-re-runs-when-i-save-the-strategy-file-and-shows-which-trades-changed-13cd

My backtest re-runs when I save the strategy file, and shows which trades changed

17+ min ago   (600+ words) I'm building a backtesting web app called Hawk Backtester on the side. This post is about one feature, a watch mode for local Python strategies. The loop that annoyed me most was edit the strategy, run it in a terminal,…...


benchlm.ai > compare > glm-5-vs-glm-5-3

GLM-5 vs GLM-5.3: Benchmarks & Cost

1+ day, 4+ hour ago   (264+ words) Updated October 10, 2026. Rank says GLM-5.3 is ahead. Price, access, and your workload can each overturn that. Public scores include evidence status and uncertainty. This is a same-family comparison, so migration details appear when the source data supports them. GLM-5.3 has…...


benchlm.ai > compare > glm-5-3-vs-glm-5-3-flash

GLM-5.3 vs GLM-5.3-Flash: Benchmarks & Cost

1+ day, 4+ hour ago   (251+ words) Updated October 10, 2026. Rank says GLM-5.3 is ahead. Price, access, and your workload can each overturn that. Public scores include evidence status and uncertainty. This is a same-family comparison, so migration details appear when the source data supports them. GLM-5.3 has…...