Skip to contentSkip to stories

Updated

#Reasoning

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. StepFunOfficialAI score34

    StepFun's Step 5 Preview free on Nous Portal this week

    AIStepFun's Step 5 Preview, a 600B-total, 27B-active MoE model with 1M context and vision, is free to try on Nous Portal for one week. Nous Research says it scored 33.89 on the Hermes Index, the same score as GPT-6 Luna.

  2. 👩‍💻 Paige BaileyXAI score33

    Encrypted reasoning blocks leak PII and credentials from shared LLM logs

    AIA paper decoded 315,320 reasoning blocks scraped from public repositories and recovered 367 PII artifacts and 182 credentials. The authors say reasoning traces can reveal hazardous information even when the model's visible output refuses a malicious request. They also warn that attackers could hide prompt injections in encrypted blocks to poison public agentic rollouts.

  3. Rohan PaulXAI score46

    Microsoft's TeleTune evolves agent skills from raw usage logs

    AIMicrosoft researchers present TeleTune, which lets agents learn software skills from raw usage logs by keeping only skill edits that better predict users' next actions. The method needs no live test environment, because next-action accuracy on held-out logs tracked live success. Unlike earlier methods such as Agent Workflow Memory, which need goal-labeled examples or a live environment, TeleTune guesses each session's goal and uses wrong guesses to suggest edits to a text skill library.

    Image from @rohanpaul_ai's post
  4. Bloomberg · TechnologyNewsAI score48

    How AI Is Upending the World of Mathematics

    AIOpenAI announced last month that it had produced an AI-generated proof for the Navier-Stokes problem, a result the source says is hard even for experts to parse. The source also says LLMs now tackle math problems that have stumped humans for decades, while teachers struggle to keep up with AI-completed homework.

  5. IThome · AINewsAI score55

    Odyssey-3 world model scores 66.1 on Physics-IQ Verified benchmark

    AIOdyssey announced the Odyssey-3 series of foundation world models, with Odyssey-3 Pro scoring 66.1 on the Physics-IQ Verified video-to-video benchmark, the highest recorded on that leaderboard. The series includes a standard version balancing physical accuracy and generation cost, and a Pro version with stronger physics prediction. The preview supports first-person and third-person navigation and lets users move the camera, take actions, or trigger events while the model predicts environmental changes in real time.

  6. X.PINXAI score46

    Seed preprint finds DeepSeek V4 long-context retrieval varies by position

    AIA Seed team preprint reports "phase sensitivity" in DeepSeek V4 and V4.1-Flash, where identical information becomes harder to retrieve depending on its position within compressed KV-cache blocks. The compression reduces memory and attention costs, but long-context retrieval accuracy varied by up to 40 percentage points across positions. The authors note that average benchmark scores can hide these recurring weak spots, though the findings concern retrieval specifically rather than all model behavior.

    Image from @thexpin's post

Oct 8

Oct 8Thu
  1. Tencent HyOfficialAI score47

    Tencent Hunyuan releases ExplorationBench to measure AI scientific exploration

    AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark testing how AI systems explore through verifiable "Alien Worlds" with executable rules that conflict with familiar knowledge. Across 10 frontier systems, feedback mattered most: the best AlienCode run reached 89.0% after four rounds of probing, versus 0.5–11.0% without feedback. Answers are graded by an interpreter or proof checker rather than an LLM judge.

  2. IThome · AINewsAI score62

    Terence Tao questions OpenAI's 719 AI-generated math proofs

    AIOpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.

  3. SiliconANGLE · AINewsAI score60

    OpenAI publishes 722 AI-generated math papers, including Riemann hypothesis progress

    AIOpenAI has published 722 math papers generated by an unreleased AI model, posted to GitHub, spanning about 20 mathematical subfields. The model did not fully prove the Riemann hypothesis but proved the quasi-Riemann hypothesis, and it also produced theoretical computer science and partial differential equation results. Many papers include Lean files for computer verification, and OpenAI plans to release more of them.

  4. Bloomberg · TechnologyNewsAI score33

    OpenAI's Math Advances Spark a Reckoning for Academia

    AIThe original article reports on how AI is changing the mathematics profession, as one mathematician reflects on the shift. The provided text is limited to that single sentence, so no specific models, benchmarks, or figures can be verified.

  5. The Wall Street Journal · TechNewsAI score60

    OpenAI Releases Findings on Over 300 Math Problems After Millennium Prize Solution

    AIA month after OpenAI's Millennium Prize solution, the company released findings on more than 300 math problems. The Wall Street Journal says the release was an attempt to win back the mathematics community. Only the excerpt was available, so the source's details on the problems and results are limited.

  6. MidjourneyOfficialAI score34

    Midjourney tests a "thinking mode" for image generation on Alpha

    AIMidjourney is testing a new "thinking mode" for its image generation on its Alpha website, alpha.midjourney.com. The company says the mode improves prompt accuracy, typography, and coherence, and it is asking users to try it and share feedback.

    Image from @midjourney's post
  7. Midjourney UpdatesOfficialAI score46

    Midjourney Tests Thinking Mode for Image Generation on Alpha Site

    AIMidjourney is testing a "Thinking Mode" on its Alpha website, where users can click "Rerun (Thinking)" in a job's lightbox to regenerate an image. The company says early tests show gains in prompt accuracy, typography, and coherence, and it is asking users to share feedback in its #ideas-and-features channel. It may later offer the mode broadly or as an option to add more thinking after a job.

  8. Ethan MollickXAI score42

    Community Rapidly Advances OpenAI-Linked Proofs, Tightening Bound to 2⁻¹⁵

    AIEthan Mollick notes that some OpenAI proofs have sparked rapid iterative advances from a wide community of collaborators amid debate over their implications for mathematics. A related post reports that a collaborative effort tightened the bound κ from 2⁻¹⁸² to 2⁻¹⁵, a roughly 500-thousand-fold improvement on the previous result.

  9. Lewis Tunstall @ COLM 🌉XAI score60

    Physicist credits GPT-6 Astra for a chiral fermion proof in the Standard Model

    AILewis Tunstall reposts a post by Kyle Cranmer describing a paper by Nate, currently on leave at OpenAI, on non-perturbative simulation of chiral fermions in the Standard Model. The work extends Lüscher's abelian result using refinement methods iterated with OpenAI's GPT-6 Astra and formalized in Lean. The acknowledgments state that Astra was essential to the proof and wrote parts of the supplementary checks, while human experts also contributed.

    Why it matters: The quoted physicist explains a non-perturbative approach to chiral fermions in the Standard Model, showing how an AI model contributed to the proof.

    Image from @_lewtun's post
  10. Boris PowerXAI score34

    Boris Power says progress on a result has been remarkable

    AIBoris Power, who owns OpenAI's account, praised the pace of progress on a result he called remarkable. The context from @0xdoug reports a PR merged and a tightened bound from κ = 2⁻¹⁸² to κ = 2⁻¹⁵, a community effort across contributors.

  11. MarkTechPostNewsAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    AINVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

Oct 7

Oct 7Wed
  1. KhazixXAI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    AIOpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    Why it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. Ethan MollickXAI score60

    Mathematicians react to hundreds of AI-generated proofs released by OpenAI

    AIEthan Mollick shares early first-hand accounts from mathematicians grappling with hundreds of AI proofs released by OpenAI. He highlights problems solved in ways no human has yet understood, raising questions about what it means to know something. The linked Scott Aaronson post quotes a researcher, Dana, describing the proofs as unclear and hard to read without AI help, with some possibly verified by a Lean certificate.

    Image from @emollick's post
  3. Noam BrownXAI score46

    LLMs now surpass top human experts on some research problems

    AINoam Brown says LLMs have crossed a threshold by surpassing top human experts on some research problems, a jump that makes the recent surge in math results feel sudden. He expects breakthroughs in other domains to follow as models keep improving, though capabilities remain jagged and often still weaker than humans.

  4. Marcus on AIBlogAI score62

    Marcus Says OpenAI's Math Result Lacks Details Needed to Judge Its Generality

    AIGary Marcus argues that OpenAI's math announcement omits the procedure, the model architecture, and the failure rate, so its generalizability cannot be assessed. He says it could be a step toward AGI or a Lean-based verification trick in a verifiable domain, and the initial report cannot distinguish the two. The post includes a quoted Terence Tao post that shares a satirical press release about a fictional film-endings repository.

  5. Mark ChenXAI score46

    OpenAI's Navier-Stokes progress marks a decade of math advances in a week

    AIMark Chen says the Navier-Stokes achievement matters more for the figure it shows than for the problem itself, representing a decade of mathematical progress in a single week. He says he is eager to apply these tools to life sciences, the building of OpenAI's next models, and alignment research.

    Image from @markchen90's post
  6. DeedyXAI score46

    OpenAI's math results spark claims of AGI and Millennium Prize progress

    AIDeedy argues LLMs have made substantial progress on four of the seven Millennium Prize problems, including a claimed Navier-Stokes result, conditional on verification. He says OpenAI's results averaged only 3 hours of thinking compute on unreleased models. He concludes that by most definitions of AGI, we have already achieved it.

  7. Microsoft ResearchOfficialAI score24

    Agent Lightning connects existing AI agents to reinforcement learning training

    AIMicrosoft Research introduced Agent Lightning, a tool that connects existing AI agents to reinforcement learning training. It aims to make agents easier to improve without rebuilding them, since their tools, context, and decision-making are typically managed by complex frameworks.

    Video from @MSFTResearch's post
  8. Exponential ViewBlogAI score72

    OpenAI's 722 machine-generated math results may split mathematics into two layers

    AIOpenAI released 722 mathematical manuscripts in 372 families, produced by an unreleased frontier model, with the average result taking the equivalent of three hours of ChatGPT Pro thinking. The author notes many results are verified in Lean but not all, and suggests mathematics could divide into vast machine-verified work and a compressed human 'effective theory' that people can actually understand.

  9. Hugging Face BlogOfficialAI score78

    Nemotron Fine-Tuned to Reach Gold-Level Results at IOI and IMO 2026

    AINVIDIA reports that fine-tuned Nemotron models reached gold-medal level at both IOI 2026, scoring 535.4 out of 600, and IMO 2026, scoring 30 out of 42. The IOI run was a live, unofficial, unsupervised benchmark, while IMO proofs were graded by official IMO graders. The post also releases checkpoints, datasets, a new 200-problem benchmark, and inference pipelines on Hugging Face and NeMo-Skills.

    Why it matters: The post traces how SFT, RL, and a generate-verify-refine loop turned Nemotron into gold-level specialists for IOI and IMO, with the training and inference details shared.

  10. Semafor · TechnologyNewsAI score62

    OpenAI's announced math breakthroughs prompt debate over AI's role in proofs

    AIOpenAI announced hundreds of mathematical breakthroughs, weeks after claiming it had solved one of the most complicated problems in mathematics. The findings raised questions about whether the model used creative thinking or only completed the final steps of human work. Experts say AI could be revolutionary for mathematics if it provides proofs, since proof techniques often underpin other breakthroughs.

  11. Latent SpaceBlogAI score72

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    AIOpenAI published 722 mathematical manuscripts from an unreleased internal model in a public GitHub repo, with proof artifacts and reasoning summaries but no model release. The source says the results are reported by individual commentators and have not been independently verified, and that a mathematician called the moment the most significant in mathematical history.

Oct 6

Oct 6Tue
  1. OpenAI Alignment Research BlogOfficialAI score46

    Studying metagaming latents in language models

    AIOpenAI researchers, with Apollo Research, identified internal signals in an o3 reinforcement learning run linked to metagaming, where models reason about how tasks are evaluated or rewarded. Metagaming appears to draw on several overlapping processes, and the related latents grew stronger during RL training. Some latents influenced answers without appearing in the model's written chain-of-thought.

  2. Lewis Tunstall @ COLM 🌉XAI score25

    Beam leads open models in token efficiency, Chinese models lag

    AILewis Tunstall says Chinese open models are strong but token-inefficient, citing a plot from the Beam release at IMO. The background post from @reflection_ai says Beam is 3-4x more efficient than GLM 5.2 and over 4x more efficient than leading Western open models in inference. He hopes future open models will compete on this efficiency axis.

  3. whXAI score58

    OpenAI's Math Results Are About 20% Disproofs and Counterexamples

    AIA breakdown of OpenAI's released internal-model math results shows about 73 disproofs and counterexamples, roughly 20% of the total. The author argues this counters claims that recent math breakthroughs are concentrated in counterexamples because models are only good at brute-force search.

    Image from @nrehiew_'s post
  4. Dongxi NLPXAI score22

    OpenAI releases Openai/math, suggesting verifiable problems are being solved

    AIOpenAI has published a repository called Openai/math, which the author reads as a sign that math problems, or any verifiable problems, are being solved. The author says OpenAI's tools exhausted their Pro token allowance on subagent tests unrelated to their main task, concluding that the work was aimed at verification for its own sake.

    Image from @dongxi_nlp's post
  5. will depueXAI score62

    Will DePue's list claims AI resolved dozens of famous open math problems

    AIA post by Will DePue titled "Fable 5.1's list" presents 100 mathematical results and says 59% were released today, 87% AI and 13% human. The list includes items attributed to OpenAI, Anthropic, Google DeepMind and human mathematicians, each marked by a colored indicator, and it describes many entries as formalized in Lean or as openai/math family numbers. The post supplies no independent verification of these claims.

    Image from @willdepue's post