Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 7

Oct 7Wed
  1. O'Reilly RadarAI score42

    Build Your Own Post-Training Pipeline: SFT, Reward Model, and PPO

    AIThe final post in O'Reilly Radar's four-part post-training series walks readers through implementing the classic ChatGPT pipeline on Qwen2.5-1.5B, covering SFT, reward model training, and PPO. The walkthrough uses torchtune for SFT and verl, a Ray-based RL framework from ByteDance's team, for reinforcement learning. The author says the goal is hands-on understanding rather than reproducing InstructGPT, which took a large team and thousands of GPU-hours.

  2. Semafor · TechnologyAI score62

    OpenAI's announced math breakthroughs prompt debate over AI's role in proofs

    AIOpenAI announced hundreds of mathematical breakthroughs, weeks after claiming it had solved one of the most complicated problems in mathematics. The findings raised questions about whether the model used creative thinking or only completed the final steps of human work. Experts say AI could be revolutionary for mathematics if it provides proofs, since proof techniques often underpin other breakthroughs.

  3. ChinaTalkAI score58

    Why an FCC ban on Chinese optical transceivers would not reduce U.S. dependence

    AIThe FCC's proposed ban on new Chinese optical transceivers targets the top of the supply stack, but the author argues it leaves the dependencies that matter untouched. The analysis traces the module, laser, indium phosphide wafer, and indium metal layers, finding that China controls the wafers and refined indium while U.S. firms depend on Chinese-made substrates. The author concludes that a module-level rule would take years to replace lost capacity and would not change control of the lower layers.

  4. Rest of WorldAI score62

    Red Sea conflict pushes Google and Meta to shift traffic onto Iraq land route

    AIGoogle and Meta have started sending some live traffic through a land route across Iraq that they had previously held in reserve, according to a person familiar with the deal. Most data between Europe and Asia still flows through subsea cables under the Red Sea, where Yemen's side of the strait is now contested and cable repairs could take months.

  5. Gizmodo · AIAI score46

    Michael Smith Sentenced to 18 Months for AI Music Streaming Royalty Fraud

    AIMichael Smith, a 54-year-old from North Carolina, was sentenced to 18 months in prison after pleading guilty to streaming-royalty fraud using automated bots and AI-generated songs. The U.S. Justice Department also ordered him to forfeit the $8,091,843.64 he received in royalties. Prosecutors said the scheme diverted royalty payments from genuine artists.

  6. Testing CatalogAI score47

    Daily AI brief covers Mistral Large 4, Google, OpenAI, and Anthropic updates

    AIMistral released Mistral Large 4 "Le Chonk", a 1T-parameter (49B active) multimodal model, with open weights planned in about three weeks. Google rolled out Nano Banana 2.1 across Gemini, AI Studio, and the Gemini API, and released EmbeddingGemma 2, a 740M-parameter open multimodal embedding model under Apache 2.0. OpenAI launched the Decisions API in beta with gpt-6-luna, returning typed answers 10x faster than the Responses API.

  7. Wired · AIAI score40

    Books by People Launches "Organic" Mark Certifying Human-Written Books

    AIBooks by People, a British startup, has received UK Intellectual Property Office approval to distribute a mark certifying that a published book was written by a human, displayed on the cover as a thumbprint-style stamp. The startup verifies manuscripts through an authorship declaration, proprietary software, and review of drafts and research materials, and five publishers have signed up so far. The mark is expected to appear on its first book by the end of the year.

  8. The Register · AIAI score38

    COSMIC bans AI-generated contributions as GNOME debates accepting AI bug reports

    AISystem76's COSMIC desktop now requires contributors to declare no LLM-generated content in pull requests, including code, comments, and descriptions. GNOME Calendar and GNOME Extensions also restrict AI-generated contributions, while GNOME developer Michael Catanzaro argues the project should accept AI-generated bug reports. Catanzaro's case rests on memory-unsafe languages such as C, C++, and Vala, and he has shortened GNOME Security's disclosure deadline from 90 days to 30, effective August 1.

  9. Ai2 (Allen Institute for AI)AI score57

    Ai2's Bolmo byte-level language models are published in Nature

    AIAi2 has published its Bolmo byte-level language model research in Nature and released new checkpoints on Hugging Face. The byteifying process converts an existing subword model into a byte-level one with a relatively short additional training run, and the paper reports that it also works for Qwen 3 8B and Llama 3 8B, producing Bwen 8B and Blama 8B. Ai2 also released Stage 1 checkpoints for researchers extending the architecture.

  10. MarkTechPostAI score58

    Meta open-sources Rebalancer, a C++ assignment solver for placement problems

    AIMeta has open-sourced Rebalancer, a C++ library with a Python interface for solving assignment problems under constraints and objectives, released under Apache 2.0. The article reports that Meta has used it for resource allocation for over 9 years and runs about 40 million problems a day, with P99 solve time of 12 seconds on 265k objects and 3.2k bins. The package can be installed with pip install rebalancer, though PyPI still classifies it as Alpha.

  11. indigoAI score60

    Meta and Sierra Announce Personal Agent Protocol for Agent-Business Interaction

    AIMeta and Sierra announced the Personal Agent Protocol, an open standard for how personal AI agents find and transact with businesses on a user's behalf. The author says it defines discovery, OAuth-based sessions, and a choice among website, API, or company agent routes, and distinguishes it from MCP, which connects agents to tools and data, and A2A, which hands tasks to another agent.

  12. Latent SpaceAI score72

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    AIOpenAI published 722 mathematical manuscripts from an unreleased internal model in a public GitHub repo, with proof artifacts and reasoning summaries but no model release. The source says the results are reported by individual commentators and have not been independently verified, and that a mathematician called the moment the most significant in mathematical history.

  13. LangChain BlogAI score42

    Deep Agents Adds Tool Binding, Pinned Skills, and Skill Reloading

    AILangChain revamped skills support in Deep Agents with three changes: tools bound to a skill load only when the agent reads that skill, pinned skills are loaded before the next model call when a user requests them, and long-running threads can pick up new or changed skills without restarting. Each skill is a folder with a SKILL.md file, and only its name and description are in context until the agent reads the full instructions.

  14. Claude BlogAI score70

    Anthropic releases Claude Haiku 5.5, its cheapest and fastest small model

    AIAnthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model. It costs around 75% less to run than Haiku 4.5 and is aimed at high-volume, cost-sensitive tasks such as summaries and classification. The release also cuts Sonnet 5.5 cache read prices by 50%, and the model is available on AWS, Google Cloud, and Microsoft Azure.

  15. Artificial Analysis ArticlesAI score60

    Anthropic releases Claude Haiku 5.5, scoring 43 on the Intelligence Index

    AIAnthropic released Claude Haiku 5.5, which scores 43 on the Artificial Analysis Intelligence Index, up 26 points from the last Haiku release. Pricing is $0.10/$0.50 per 1M input/output tokens up to 100k tokens, rising to $0.50/$2.50 above that, but at max effort it uses about 162k output tokens per Intelligence Index task, roughly 3x GPT-6 Luna.

    Why it matters: The benchmark shows Haiku 5.5 scores well but uses far more output tokens than GPT-6 Luna, so cost per task matters beyond list price.

  16. Mastra BlogAI score60

    Mastra Connect adds ready-made tools for services like Linear and Notion

    AIMastra Connect is a public beta that lets Mastra projects connect providers such as Linear, Notion, and Slack, giving agents and workflows ready-made tools. Connect launches with 23 providers, almost 900 tools, and 7 hosted MCP providers, and it is free to use on Mastra platform during beta. Developers can add connections via the CLI or dashboard, limit tools with glob filters, and call a provider's SDK directly with credential() when a tool is missing.

    Why it matters: The post shows how connected services become agent tools, and how credentials and access limits are managed, which is useful for building agent workflows.

  17. Claude BlogAI score66

    Claude skill commands build evals and hillclimb them against overfitting

    AIAnthropic added build-eval and hillclimb commands to its claude-api skill for designing evaluations and iteratively improving applications against them. The article covers eval design principles, including production-representative tasks, headroom and low variance, and guards against overfitting through train/test splits. Two examples report results: a customer support benchmark where cost fell to under half while accuracy rose, and a claude-api skill eval that rose from 66% to 88%.

    Why it matters: The article gives a concrete workflow for designing evals and hillclimbing without overfitting, with two worked cost and performance examples that show the tradeoffs.

  18. LangChain BlogAI score63

    Managed Deep Agents v0.9 adds agent schedules, per-run configuration, and Slack reactions

    AILangChain released Managed Deep Agents v0.9 in Public Beta, adding a Schedules SDK, per-run agent configuration, and Slack reactions. Agents can create reminders, follow-ups, and recurring tasks mid-conversation, running as the requesting user and posting results back to the originating channel. Per-run configuration lets one deployment choose the model, instructions, skills, MCP servers, and sandbox based on the run's context, and Slack reactions are on by default with a 👀 emoji.

    Why it matters: The release shows how one agent deployment can be configured per run by channel or repo, separating tool access from model instructions.

Oct 6

Oct 6Tue
  1. OpenAI Alignment Research BlogAI score46

    Studying metagaming latents in language models

    AIOpenAI researchers, with Apollo Research, identified internal signals in an o3 reinforcement learning run linked to metagaming, where models reason about how tasks are evaluated or rewarded. Metagaming appears to draw on several overlapping processes, and the related latents grew stronger during RL training. Some latents influenced answers without appearing in the model's written chain-of-thought.

  2. meng shaoAI score35

    Claude Code's html-plan plugin turns plans into reviewable HTML pages

    AIClaude Code developer Thariq (@trq212) released html-plan, a plugin that makes Claude Code generate self-contained single-file HTML plans instead of lengthy Markdown. The page organizes the plan into a layered tree with progressive disclosure, numbered decision points, and in-page feedback that can be pasted back into Claude Code. Install it with claude plugin marketplace add anthropics/claude-plugins-community, then claude plugin install html-plan@claude-community.