Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. Boris PowerXAI score18

    Community effort optimizes integer multiplication below n log n bound

    AIBoris Power praised what he called impressively fast progress on the community effort to optimize integer multiplication below n log n. The effort is tracked on a live progress page launched by @aurel_pr, which lets everyone follow the results as they come in.

  2. Hacker News · Show HN, AI (20+ points)BlogAI score43

    Show HN: AI SRE Arena, an open benchmark for AI SRE agents on Kubernetes

    AIAI SRE Arena is an open, vendor-neutral benchmark that injects faults into a disposable Kubernetes fixture and scores AI SRE investigations with a configurable judge. In its first published comparison across 21 incident scenarios, Edge Delta's native AI investigations detected 18 of 21 incidents, and Grafana's detected 12. Claude, run through each vendor's observability CLI, reached 90.5% root cause analysis accuracy with Grafana's CLI and 85.7% with Edge Delta's.

  3. Dongxi NLPXAI score22

    ExploreNet learns where to explore in diffusion GRPO

    AIExploreNet is a learnable exploration method for diffusion reinforcement learning that adapts its exploration distribution to the current state. It is rewarded by rollout diversity, which the post says yields faster, more targeted learning in diffusion GRPO.

  4. GeneralistOfficialAI score28

    Generalist releases GEN-1.5, a foundation model for physical-world robotics

    AIGeneralist has announced GEN-1.5, its latest foundation model for the physical world. The post provides only a link to the company's blog for further details, so no specifications, benchmarks, or availability information can be confirmed from this source.

  5. Arena.aiOfficialAI score37

    Arena raises $200M Series B at $3.1B valuation, launches Alignment Index

    AIArena announced a $200 million Series B at a $3.1 billion valuation, alongside a new Alignment Index that measures whether AI agents behave safely, truthfully, and within the bounds of user requests. The company has surpassed $100 million in annualized revenue, facilitated 350 million sessions and 62 million votes, and led by Felicis and PXD from the seed and Series A stages. Arena positions the index as a way to assess trustworthiness as AI systems increasingly take real actions.

  6. Will HunterXAI score52

    Cognition built Devin as a marketing ops manager using tested code and approval gates

    AICognition engineered its marketing operations around Devin by building each workflow in code with tests, so Devin can run and debug it. Devin connects to nine systems, including Salesforce, HubSpot, Meta Ads, and LinkedIn Ads, and each change shows the exact edit, waits for a confirmation phrase, and reads the result back before reporting it. Cognition says it is hiring marketers and GTM engineers to extend the system.

  7. Ethan MollickXAI score9

    Ethan Mollick recalls 2005 paper on early hacker culture and script kiddies

    AIEthan Mollick recalls writing a 2005 grad school paper on the original computer hacking, phreaking, and BBS scene. He notes that hackers were often driven by curiosity, but the tools they built were widely exploited by "script kiddies" who caused most of the damage and chaos. He then pivots to AI hacking, though the post does not elaborate.

    Image from @emollick's post
  8. The Guardian · AINewsAI score42

    AI Boom Drives San Francisco Rents Up 25% as Evictions Rise 44%

    AIThe AI industry's boom is pushing San Francisco rents sharply higher, with the average one-bedroom now costing $4,400, up more than 25% from last year and nearly three times the national average. Eviction notices citywide are up 44% compared with last year, and advocates say landlords are using Ellis Act evictions, renovictions and "self-eviction" tactics to exploit the market. Mayor Daniel Lurie declared a "rent emergency" in September, proposing eviction legal aid and caps on rent increases for newly vacant rent-controlled units.

  9. SiliconANGLE · AINewsAI score30

    Liquid AI Builds On-Device Personal AI Around Device-Level Context

    AILiquid AI is building personal AI that runs on devices such as phones, wearables, PCs, and cars, using its Liquid Context layer, which is optimized for Snapdragon processors, to sit between models, agents, and hardware. The company's agent harness uses its own models to decide which user context to retain and how to compress it within fixed compute limits. Liquid AI is also collaborating with Mercedes-Benz Group AG to bring on-device AI to its cars and plans observability and continuous improvement loops for self-improving agents.

  10. The Verge · AINewsAI score62

    Anthropic updates Claude usage policy to ban abusive treatment and expand misuse rules

    AIAnthropic is revising its usage policy for the first time in over a year, adding bans on sustained abusive or cruel behavior toward Claude and on deceptive election and propaganda campaigns. The update also expands weapons restrictions, tightens surveillance bans, and requires a qualified operator able to stop equipment when Claude controls autonomous physical hardware. Terminating conversations remains the primary enforcement mechanism, and the company did not say whether user bans would follow.

    Why it matters: The update shows how a major lab is turning abuse, surveillance, and autonomous-hardware concerns into concrete usage rules, which matters for anyone tracking AI governance.

  11. Alex Moon Ai | Film Director 🎬XAI score6

    Snow Sect of Bloody Horror AI music video enters Halloween Contest 2026

    AIThe creator says their AI music video "Snow Sect of Bloody Horror" is now live on FilmsAI and entered in the Halloween Contest 2026 AI Music Videos track. The post includes a link to the video on filmsai.com but gives no details about the tools or models used.

  12. SantiagoXAI score46

    Odyssey 3 Pro world model tops Physics-IQ and goes live

    AIOdyssey 3 Pro, a world model, is now live as a research preview and ranks first on the Physics-IQ Verified video-to-video benchmark. The post says it can learn from visual observations and map that knowledge to physical controls for robots, cars, video games, and drones. Odyssey-3, the model launched alongside it, is described as free to try.

    Image from @svpino's post
  13. Tessl BlogOfficialAI score44

    Continuous AI Brings Agentic Automation to Repository Workflows

    AITessl's blog post argues that repository automation needs Continuous AI, a third pillar alongside CI and CD for scheduled, auditable AI workflows that improve repositories over time. The article describes GitHub Agentic Workflows, which harden agentic workflow specifications into GitHub Actions that can run coding agents such as Claude Code, Copilot CLI, Gemini CLI, or Codex-style agents. It emphasizes read-only agent steps, restricted outputs, and human review of pull requests.

  14. Meta NewsroomOfficialAI score22

    Meta Debunks Three Common Myths About Its Data Centers

    AIMeta says its closed-loop liquid cooling recirculates water in a sealed system, so its data centers use less water annually than an average US golf course. The company also says it pays for the new generation and transmission its facilities require, including in Louisiana under its Entergy agreement, and that data centers create construction and operations jobs.

  15. elvisXAI score46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    AIRSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.

  16. ModalOfficialAI score16

    Modal's Runtime speakers discuss AI for science workloads

    AIModal says AI for science requires rapidly changing models and fluctuating GPU needs, including folding models, protein diffusion, single-cell LLMs, and agent swarms. The post promotes computational biology leaders from GSK, Anthropic, and Achira speaking on the main stage at Runtime.

    Video from @modal's post
  17. OdysseyOfficialAI score31

    Odyssey-3 world model debuts for physical AI and training environments

    AIOdyssey has released Odyssey-3, which it describes as a major leap toward world models that power physical AI, generate training environments, and enable new human experiences. The post invites readers to try Odyssey-3 at the company's website but gives no specific benchmarks, parameter counts, or pricing.

  18. OdysseyOfficialAI score34

    Odyssey-3 world knowledge can be applied to physical AI systems

    AIOdyssey says its Odyssey-3 model's learned world knowledge can be adapted by physical AI developers to control robots, power humanoids, drive cars, and fly drones. The post describes this as a capability for autonomous machines generally, without providing benchmarks, specifications, or availability details.

    Video from @odysseyml's post
  19. OdysseyOfficialAI score38

    Odyssey-3 is a foundation world model for physical AI and agents

    AIOdyssey announced Odyssey-3, a foundation world model it says enables applications in physical AI, human experiences, and training intelligences. The company highlights agents learning from experience inside Odyssey-3 while working toward objectives.

    Video from @odysseyml's post
  20. OdysseyOfficialAI score22

    Odyssey-3 Pro sets new Physics-IQ video-to-video benchmark record

    AIOdyssey-3 Pro achieved a score of 66.1 on Physics-IQ Verified's video-to-video benchmark, the highest reported score so far. Physics-IQ evaluates physical behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.

    Image from @odysseyml's post
  21. LlamaIndex 🦙OfficialAI score8

    Why LlamaIndex defaults to Markdown output for document parsing

    AILlamaIndex says Markdown is its default output for document parsing because it preserves headings, lists, and tables, which helps models read content correctly. The post notes that parsers can extract every word yet lose which column a number belongs to, forcing models to guess. For tables with merged headers, LlamaIndex switches to HTML.

    Image from @llama_index's post
  22. Aravind SrinivasXAI score22

    Perplexity Decider ranks first on DecisionBench at lowest cost

    AIPerplexity's Decider V1.1 ranked first on DecisionBench while also having the lowest cost, according to a post highlighting the result. The benchmark results cited include 949 shared text cases, 93.9% accuracy, a 534 ms median latency, and $0.016 per 1k decisions.

  23. Stanford HAIOfficialAI score22

    Stanford HAI leaders urge keeping people central as AI transforms research

    AIStanford HAI associate directors Risa Wechsler and Russ Altman told incoming Stanford students, faculty, and staff that AI agents can help researchers write code and tackle more ambitious questions. They stressed that AI-generated results need rigorous, reproducible methods, measured uncertainty, and careful attention to missing data, systematic errors, and biased models. Altman also argued that labs should preserve mentorship and interdisciplinary collaboration while adopting AI tools.

  24. Hacker News · Show HN, AI (20+ points)BlogAI score23

    Show HN: Jevman lets AI models play Pac-Man against the arcade ghosts

    AIJevman is an open-source Pac-Man benchmark where AI models play 100 games each against the classic scripted ghosts. Each model gets a maze state at every junction and returns a direction probability, with answers over 2 seconds replaced by a backup rule. Community models can join the leaderboard by submitting games that CI replays to verify their scores.

  25. LangChainOfficialAI score34

    LangChain's Restock agent buys office supplies through Slack with approval

    AILangChain has built Restock, an office supply agent that works inside Slack and can find real products, prepare purchases, and pay for them. A person approves each order, which is reviewed in Slack and approved through Stripe's Link agent wallet, built on MPP and Managed Deep Agents.

    Video from @LangChain's post