Skip to contentSkip to stories

Updated

#Industry news

Showing low-relevance items too. Hide low-relevance items

Feb 25

Feb 25Wed
  1. Yi TayXAI score62

    Aletheia math research agent solves 6 of 10 FirstProof problems

    AIAletheia, a math research agent, autonomously solved 6 of 10 FirstProof problems without modification, the best result in the inaugural challenge. The author says this is bigger than the IMO-gold achievement from last year, and the results were evaluated by experts with best-of-2 scoring.

  2. Quoc LeXAI score53

    Google's Aletheia math agent solves 6 of 10 FirstProof problems

    AIQuoc Le announced that Aletheia, a math research agent, autonomously solved 6 of 10 FirstProof problems, the best result in the inaugural challenge. The post says this exceeds last year's IMO-gold achievement and points to a paper and thread for full details. The accompanying figure shows 10 unmodified problems, 6 candidate solutions per agent, and expert evaluation yielding 6 solved problems on a best-of-2 basis.

Feb 24

Feb 24Tue
  1. Replit BlogOfficialAI score43

    Replit Pro launches at $100/month as Core drops to $20/month

    AIReplit launched a $100/month Pro plan with Turbo Mode, pooled credits for up to 15 builders, and priority support, while cutting Core from $25 to $20 per month and letting it invite up to 5 collaborators. The Teams plan is being sunset, with Teams users automatically upgraded to Pro at no additional cost for the rest of their term. Economy and Power Modes for Agent are available on all paid plans.

Feb 23

Feb 23Mon

Feb 19

Feb 19Thu
  1. Benedict EvansBlogAI score62

    Benedict Evans questions whether OpenAI can build a durable competitive lead

    AIBenedict Evans argues that OpenAI lacks a clear competitive lead, since frontier models are close in capability and its user base shows shallow engagement. He contends that its capex-heavy platform strategy does not yet create the network effects that powered Windows or iOS, leaving execution as the main advantage.

Feb 14

Feb 14Sat
  1. Jakub PachockiXAI score22

    OpenAI now believes its #1stProof problem 2 solution is likely incorrect

    AIJakub Pachocki, OpenAI's account holder, says that after official #1stProof commentary, community analysis, and external expert review, the team now believes its solution to problem 2 is likely incorrect. He thanked reviewers for their engagement and said the team looks forward to continued review.

  2. Oriol VinyalsXAI score13

    Oriol Vinyals returns to Bay Area to continue building Gemini

    AIGoogle DeepMind researcher Oriol Vinyals announced he is moving back to California after ten years in London, to keep working on Gemini and toward AGI. The post is a personal update with no product, benchmark, or technical details.

Feb 12

Feb 12Thu

Feb 11

Feb 11Wed
  1. Yi TayXAI score67

    Aletheia math research agent produces two papers and solves open Erdős problems

    AIYi Tay introduces Aletheia, a math research agent powered by an advanced version of Gemini Deep Think. The post says it produced two publishable papers, one fully automatic and one human-AI collaboration, and solved multiple open Erdős problems. The attached image shows a Google DeepMind paper titled "Towards Autonomous Mathematics Research" with a generator, verifier, and reviser loop.

    Image from @YiTayML's post
  2. Quoc LeXAI score22

    Gemini Deep Think advances mathematical and scientific research

    AIGoogle's latest blog post describes how Gemini Deep Think is being applied to advance mathematical and scientific research. The post itself gives no specific benchmarks, figures, or release details beyond this announcement.

Feb 9

Feb 9Mon

Feb 7

Feb 7Sat

Feb 4

Feb 4Wed
  1. Guillaume Lample @ NeurIPS 2024XAI score8

    Mistral's audio team teases upcoming announcements

    AIGuillaume Lample congratulated Mistral's audio team on their work and said more would be announced soon. The post gives no model names, benchmarks, or release details.

Feb 3

Feb 3Tue

Jan 27

Jan 27Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score32

    Cognition opens London office to expand Devin autonomous coding for European businesses

    AICognition is opening a London office to expand rollout of Devin, its autonomous software engineering agent, to leading European businesses. The company says finance has emerged as a clear use case, with Goldman Sachs, Santander, Citi, and BNY among partners using Devin for modernization, migration, security remediation, and codebase documentation.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score38

    Cognizant Partners with Cognition to Scale Devin and Windsurf Across Its Engineering Teams

    AICognizant is deploying Cognition's Devin autonomous software engineer and Windsurf agentic IDE across its engineering organization and global client base. Engineers already use Windsurf for agent-assisted coding and are exploring Devin for end-to-end tasks such as code migration, refactoring, testing, and maintenance. Cognition will embed forward-deployed AI engineers to support project selection, engineer enablement, and ROI measurement.

Jan 15

Jan 15Thu
  1. Guillaume Lample @ NeurIPS 2024XAI score8

    Mistral posts Master's and PhD job openings via Lever links

    AIGuillaume Lample, owner of the Mistral account, shared two job postings on Lever, one for a Master's position and one for a PhD position. The post provides only the links and no details about the roles, requirements, or research focus.

  2. BAAIOfficialAI score3

    BAAI invites visitors to AAAI 2026 booth and recruits talent

    AIBAAI will host a booth at AAAI 2026 in Singapore Expo Hall 3, Booth A47, from January 22 to 25, 9:00 to 17:00. The booth offers a free chat on large models, embodied intelligence, AI agents, and AI for Science, and BAAI is also recruiting, with applicants directed to its careers page or to email CVs with the subject "AAAI2026."

    Image from @BAAIBeijing's post

Jan 14

Jan 14Wed
  1. Ahmad Al-DahleXAI score38

    Ahmad Al-Dahle joins Airbnb as CTO after Llama open-source work

    AIFormer Meta AI leader Ahmad Al-Dahle announced he is joining Airbnb as CTO, citing Meta's open-sourcing of Llama, which has reached over 1.2 billion downloads and 60,000+ derivatives. He said the next challenge is applying advancing model capabilities to products that connect people with real places, working with Airbnb CEO Brian Chesky.

Jan 6

Jan 6Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score42

    Infosys partners with Cognition to deploy Devin AI software engineer across its enterprise

    AIInfosys will deploy Cognition's Devin, an autonomous AI software engineer, across its own teams and global client base to expand delivery capacity. The rollout begins in its Financial Services practice, covering banking, payments, capital markets, insurance, and wealth management, and is planned to extend to retail, energy, and healthcare. Over the past six months, Infosys reports material productivity gains, including COBOL and JCP servlet migrations completed in record time.

Dec 19, 2025

Dec 19, 2025Fri
  1. Nick TurleyXAI score28

    ChatGPT Pro users can gift three months of ChatGPT Plus to friends

    AIOpenAI now lets ChatGPT Pro users give friends three months of ChatGPT Plus access, though gift cards are not available. Pro users who held that status as of December 1 will find their share link in email or as a notification in ChatGPT on the web.

Dec 11, 2025

Dec 11, 2025Thu

Dec 8, 2025

Dec 8, 2025Mon
  1. ReflectionOfficialAI score14

    Reflection AI welcomes Brandon Damos from Meta Superintelligence Labs

    AIReflection AI announces that Brandon Damos has joined the company in New York City as part of its effort to build new frontier models. Damos, who previously spent six years at Meta's Fundamental AI Research lab, says he will work on post-training and reinforcement learning pipelines to advance capabilities and alignment.

Dec 4, 2025

Dec 4, 2025Thu
  1. ARC PrizeOfficialAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

  2. Quoc LeXAI score38

    Google DeepMind launches new Gemini reasoning research team in Singapore

    AIQuoc Le announced that Google's Gemini team is recruiting for a new Singapore team focused on advanced reasoning and LLM reinforcement learning. The post invites exceptional engineers and researchers to join the team, which is led by Yi Tay and reports into Quoc Le's Mountain View group.

  3. Yi TayXAI score38

    Google DeepMind's Gemini team launches new reasoning research group in Singapore

    AIYi Tay announced that Google DeepMind's Gemini team is starting a new research team in Singapore focused on advanced reasoning, LLM/RL, and improving frontier models such as Gemini and Gemini Deep Think. The team is led by Tay and reports to Quoc Le's broader team in Mountain View, which recently contributed to IMO and ICPC gold medal results with Gemini Deep Think. The team is starting small and is recruiting exceptionally capable engineers and researchers from the region and beyond.

    Image from @YiTayML's post

Dec 2, 2025

Dec 2, 2025Tue
  1. ReflectionOfficialAI score4

    Reflection AI invites visitors to meet its team at NeurIPS booth

    AIReflection AI announces that members of its research team and founders Ioannis Antonoglou and Misha Laskin will be available at its NeurIPS booth. The post is an event invitation and includes no product, model, or technical details.

    Image from @reflection_ai's post

Dec 1, 2025

Dec 1, 2025Mon
  1. ReflectionOfficialAI score12

    Reflection AI welcomes Ghorbani to lead Science of Scaling team

    AIReflection AI announced that Ghorbani has joined as leader of its Science of Scaling team. Ghorbani, who spent three years at OpenAI, says the team will deepen scientific understanding of large-scale learning and turn compute into intelligence as efficiently and predictably as possible.

Nov 26, 2025

Nov 26, 2025Wed
  1. ReflectionOfficialAI score16

    Reflection hosts NeurIPS panel on open AI ecosystems

    AIReflection says it will be at NeurIPS in San Diego next week, with its team staffing a booth and a panel on open AI ecosystems. The panel features speakers from Reflection, Meta, AI2, Berkeley's Ray, and SGLang, and will discuss what building open ecosystems requires, lessons from China, and the meaning of transparency and sovereignty.

Nov 25, 2025

Nov 25, 2025Tue
  1. Oriol VinyalsXAI score14

    Oriol Vinyals to receive honorary doctorate from UPC in Barcelona

    AIGoogle DeepMind researcher Oriol Vinyals is traveling to Barcelona to receive a Doctor Honoris Causa from his alma mater, UPC. He also announced a Thursday master class titled "From AI to AGI: The Quest for True Intelligence," with registration details linked in the post.

    Image from @OriolVinyalsML's post

Nov 19, 2025

Nov 19, 2025Wed
  1. Stability AIOfficialAI score36

    Warner Music Group and Stability AI Partner on Responsible AI Music Creation Tools

    AIWarner Music Group and Stability AI announced a collaboration to build professional-grade, ethically trained AI tools for artists, songwriters, and producers. The companies will work directly with artists to shape tools that enhance the creative process while protecting creators' rights and revenue opportunities. Stability AI's Stable Audio models, trained exclusively on licensed data, are cited as the basis for its commercially safe generative audio offerings.

Nov 18, 2025

Nov 18, 2025Tue
  1. Quoc LeXAI score38

    Gemini 3 Deep Think scores 45.1% on ARC-AGI-2

    AIGoogle's Gemini 3 Deep Think (Preview) reaches 45.14% on ARC-AGI-2's semi-private eval at $77.16 per task. ARC Prize says this doubles the prior state of the art, with Gemini 3 Pro scoring 31.11% at $0.81 per task.

Nov 13, 2025

Nov 13, 2025Thu

Nov 3, 2025

Nov 3, 2025Mon
  1. ARC PrizeOfficialAI score38

    ARC Prize Launches Verified Program to Certify ARC-AGI Benchmark Scores

    AIARC Prize Foundation announced ARC Prize Verified, a program that certifies frontier model scores on the ARC-AGI benchmark using hidden test sets and adds a third-party academic panel to audit and open-source its testing process. Five AI labs, including Google and xAI, are sponsoring ARC-AGI-3 development, and the foundation says donations do not influence verification scoring. Models that pass verification will appear on the official leaderboard with a verification badge.