Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. Arena.aiOfficialAI score36

    Claude Haiku 5.5 debuts at #30 in Code Arena WebDev

    AIClaude Haiku 5.5 (High) debuted at #30 with 1587 points in Code Arena: WebDev, just outside the Pareto frontier. It matches GPT‑6 Luna's price at $0.10/$0.50 per 1M input/output tokens while scoring six points higher. Arena says it costs 90% less than Haiku 4.5 for +257 points, and 95% less than Sonnet 5.5 (High) for -128 points.

    Image from @arena's post
  2. Vercel DevelopersOfficialAI score38

    Vercel AI Gateway adds OpenAI Ultrafast mode with GPT-6.1 Sol support

    AIVercel says OpenAI's Ultrafast mode is now available on AI Gateway, with support for GPT-6.1 Sol. Developers can enable faster output for interactive apps and coding agents by setting service_tier to ultrafast per request, including over WebSocket.

  3. AgentPhone (YC P26)XAI score16

    AgentPhone now offers UK phone numbers for AI agents

    AIAgentPhone announces that UK phone numbers are now available, letting users give their AI agent its own UK number. To celebrate, the company is offering $10 in free credits to new users who sign up at agentphone.ai and DM their signup email.

    Video from @AgentPhoneHQ's post
  4. OpenRouterOfficialAI score40

    Grok Imagine Video 1.5 Lite now available on OpenRouter

    AIOpenRouter now offers xAI's Grok Imagine Video 1.5 Lite for text-to-video and image-to-video generation. The quoted post from Grok Imagine lists pricing of $0.02 per second at 480p, $0.03 per second at 720p, and $0.14 per second at 1080p.

  5. Claude Code · GitHub ReleasesOfficialAI score56

    Claude Code v2.1.295 adds hook failure blocking and gateway controls

    AIClaude Code v2.1.295 adds onFailure: "block" for command and HTTP hooks, so a hook that cannot start, times out, or exits unexpectedly blocks the action. The release also adds an optional models list for Claude apps gateway upstreams, plus upstream_request_id in the inference audit event, and fixes a range of MCP, plugin, and terminal issues.

  6. @timnitGebru (@dair-community.social/bsky.social)XAI score19

    Filmmaker says AI documentary platforms "eugenist cult leaders" and promotes a rival film

    AIA commentator says they regret taking part in "The AI Doc," arguing it platforms eugenist cult leaders such as Eliezer Yudkowsky who lack expertise on AI. They recommend watching the documentary "Ghosts in the Machine" instead and point readers to Emily M. Bender's review comparing the two films.

  7. Allie K. MillerXAI score7

    Allie K. Miller jokes about "EEO," optimizing email for AI engines

    AIAllie K. Miller jokes that people obsess over web-based GEO while neglecting "EEO," or email engine optimization. She says she is drafting side notes to AI tools left and right, playfully suggesting email is a channel for influencing AI outputs.

    Image from @alliekmiller's post
  8. LiveKitOfficialAI score4

    LiveKit and Modal host AI Agents Speakeasy in LA on October 14

    AILiveKit is hosting an AI Agents Speakeasy with Modal during LA Tech Week on Wednesday, October 14, from 6 to 9 p.m. in Los Angeles. The event offers cocktails, food, and conversation with people building AI agents, with RSVPs available through a Partiful link.

    Image from @livekit's post
  9. Vercel DevelopersOfficialAI score26

    FLUX 3 Image is now available on Vercel AI Gateway

    AIVercel says FLUX 3 Image, from Black Forest Labs, is live on AI Gateway for generating and editing images. It supports up to 10 reference images and 4K output.

    Image from @vercel_dev's post
  10. Eugene SmartsXAI score44

    Grok Bot runs named AI coworkers on one shared persistent cloud computer

    AIGrok Bot, from dot.com, lets an office roster of named AI workers such as Chief, Sales Outbound, Talent Scout, and Inbox Manager share one persistent cloud computer. Sales Outbound uses Hex and Salesforce to queue 36 personalized outreach drafts overnight, with human review before anything is sent. Isolation is set per user rather than per bot, so every worker shares the same browser cookies, files, and authenticated SaaS sessions.

    Image from @EugeneSmarts's post
  11. Artificial AnalysisOfficialAI score10

    Artificial Analysis publishes its Cyber Index evaluation results

    AIArtificial Analysis has released full evaluation results for its Cyber Index, which the post links at artificialanalysis.ai. The post itself gives no models, scores, or other figures, so the details are only available through the linked page.

  12. Artificial AnalysisOfficialAI score22

    Artificial Analysis launches Cyber Index Alliance with IBM and NVIDIA

    AIArtificial Analysis has formed the Cyber Index Alliance to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. Current members are Collinear, IBM, NVIDIA, and Vercel, and partners contribute expert input on the Index design and implementation, plus datasets and external research. Organizations interested in joining can contact cyber@artificialanalysis.ai.

    Image from @ArtificialAnlys's post
  13. Artificial AnalysisOfficialAI score38

    GPT-6 Sol (Daybreak Blue) tops Artificial Analysis Cyber Index

    AIGPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model and ranks #1 on the Index. Compared with the publicly available GPT-6 Sol, it shows its largest gains on CyberGym-E2E, the benchmark where the most safety refusals are observed.

    Image from @ArtificialAnlys's post
  14. Artificial AnalysisOfficialAI score46

    GPT-6 Sol tops Cyber Index at $1.77 per task

    AIGPT-6 Sol (Daybreak Blue, max) ranks first on the Artificial Analysis Cyber Index at a Cost per Task of $1.77. That is significantly cheaper than other leading models, including Grok 4.7 (xhigh) at $11.67 per task.

    Image from @ArtificialAnlys's post
  15. Artificial AnalysisOfficialAI score62

    GPT-6 Sol (Daybreak Blue) leads Artificial Analysis Cyber Index with trusted access

    AIArtificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now leads the leaderboard. The model is available only through OpenAI's Daybreak program and records no safety blocks, improving 32 points over the publicly available GPT-6 Sol (max). It costs $1.77 per task, below Grok 4.7 (xhigh) at $11.67 per task.

    Image from @ArtificialAnlys's post

    This story has a top pick“GPT-6 Sol Daybreak Blue leads the Artificial Analysis Cyber Index”

  16. Boris PowerXAI score38

    OpenAI releases GPT-6.1 Sol ultrafast with instant steering improvements

    AIOpenAI has improved steering so the model reacts much faster to user adjustments, letting people course-correct in real time. The company is also releasing GPT-6.1 Sol ultrafast, which it says works very well alongside the steering improvement.

  17. The Guardian · AINewsAI score46

    Teen hiker rescued after Claude's route led him to a rock wall in B.C.

    AIBryce Vincent Gowryluk, 16, had to be rescued by North Shore Rescue after following route directions from the AI chatbot Claude that took him to the base of the Widowmaker Arete, a 1,700-foot wall in British Columbia. He had asked Claude to plan a route from Grouse Mountain to Crown Mountain and back, and he was ill-equipped for the rock climbing the wall requires. Rescue manager Paul Markey warned hikers not to rely blindly on AI for route planning.

  18. Epoch AI · The Epoch BriefOfficialAI score49

    Epoch AI estimates AI agent counts and falling model costs in its October 2026 brief

    AIEpoch AI estimates AI chips shipped through 2027 could support about 30 to 170 million concurrent frontier-model agents, or nearly 2 billion with cheaper models. Researchers David Roodman and Luke Emberson find the cost of a fixed level of AI performance has fallen about 47% per quarter over the past three years, around 13× per year.

  19. Wired · AINewsAI score22

    Designer Jessica Hische Faces Backlash Over Meta's Muse AI Logo Work

    AILettering artist Jessica Hische designed the M logo for Meta's new AI assistant Muse, a two-day freelance sprint she says paid about $12,500. She said she took the job to cover payroll for her Oakland shops, and the decision drew heavy criticism from fellow designers on Threads.

  20. OpenRouterOfficialAI score43

    Sol Ultrafast for GPT-6.1 Sol is now live on OpenRouter

    AIOpenRouter now offers Sol Ultrafast, a version of GPT-6.1 Sol, available through its platform. According to OpenAI Devs, the Ultrafast mode delivers near-Astra intelligence at up to 8x the speed of Sol Standard and is rolling out in the API, Codex, and ChatGPT Work.

  21. Stability AIOfficialAI score30

    Stability AI's SemanTok makes video world models more efficient

    AIStability AI's Interactive Research team introduced SemanTok, which makes early video tokens more semantically meaningful so the representation is easier to predict. According to the post, a model using SemanTok matches or beats the performance of a model more than three times its size. The approach targets more efficient autoregressive video generation.

    Video from @StabilityAI's post
  22. Cassandra UnchainedXAI score23

    Michael Burry argues AI cloud oligopoly spending faces financing risks

    AIMichael Burry argues that Microsoft, Amazon, Google, and Meta are spending on AI to expand an oligopoly, with OpenAI, Anthropic, and Oracle seeking a share. He says data center financing tied to private equity, private credit, and insurers is showing strain, with long-term rates rising before the buildout timelines play out.

  23. elvisXAI score48

    Google's FlowAgent auto-repairs failing tests inside code review

    AIGoogle proposed FlowAgent, a ReAct-style agent that generates and validates fixes for pre-submit test failures and shows them in its code review tools. Two abstention filters, before and after execution, suppress weak suggestions; in a manual review of 195 real failures, 67.18% of fixes were correct. After the Google-wide launch, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554.

    Image from @omarsar0's post
  24. Mustafa SuleymanXAI score34

    Suleyman posts emoji reaction to Anthropic's Claude abuse policy

    AIMustafa Suleyman, Microsoft AI chief, responded to Anthropic's policy change with only an emoji post, offering no detailed comment. The quoted background post says abusive behavior toward Claude will violate Anthropic's Usage Policy effective November 12, 2026.

  25. Sherwin WuXAI score60

    Artificial Analysis adds a hallucination gate to the Harvey LAB legal benchmark

    AIArtificial Analysis, working with Harvey, released LAB-AA v1.1, which credits a legal task only when every rubric criterion passes and no material hallucination appears. Grok 4.7 (xhigh) leads at 9.4%, and more than 60% of otherwise passing results contain a material hallucination. Sherwin Wu, an OpenAI-affiliated account, reposted the announcement and said the original LAB results were puzzling and that GPT-6 Astra had one of the lowest hallucination rates.

    Why it matters: The post shows how adding a hallucination gate to a legal agent benchmark reorders the leaderboard, which matters for judging models in legal work.

  26. The DecoderNewsAI score65

    Anthropic launches Claude Dashboards and Motion features in beta

    AIAnthropic launched two beta features for Claude: Dashboards, which turns connected data sources like BigQuery, Databricks, Snowflake, or Salesforce into auto-updating live dashboards from text prompts, and Motion, which creates animated explainer videos from text, diagrams, and images. Dashboards is available to paid users and Motion to Team and Enterprise plans, while Docs, Slides, and Design leave beta and work across all plans, including free accounts.

  27. LumaOfficialAI score10

    Luma Partners with Claude on Motion Launch Initiative

    AILuma announced a partnership with Claude, linked to a launch called Motion, with details available on its website. The post itself provides no further specifics about the features, terms, or outcomes of the collaboration.

  28. LumaOfficialAI score25

    Luma lets users continue Claude Motion animations in Luma

    AILuma says users can bring Claude Motion animations into Luma to resize them for different formats and refine them for shipping. Claude Motion is in beta on Claude Team and Enterprise plans.

    Image from @LumaLabsAI's post
  29. Alex HeathXAI score38

    Qualcomm CEO Cristiano Amon on AI phones, glasses, and 6G

    AIQualcomm CEO Cristiano Amon discusses the coming AI smartphone supercycle, arguing phones will not disappear as agents use personal context. He also expects smart glasses to become the largest AI wearable category, and covers Qualcomm's Modular acquisition as an alternative to Nvidia's CUDA software and its data center strategy. The conversation, recorded live at the Snapdragon Summit in Hawaii, also covers 6G being designed for AI.

    Video from @alexeheath's post
  30. TiboXAI score62

    OpenAI rolls out GPT-6.1 Sol ultrafast with faster steering

    AITibo, an OpenAI team member, says GPT-6.1 Sol ultrafast is rolling out today in the API, Codex, and ChatGPT Work. He says it offers near-Astra intelligence at up to 8x the speed of Sol Standard. The post also says improved steering now lets the model react faster to user adjustments in real time.

    Why it matters: The post specifies the new Ultrafast option, its availability across API, Codex, and ChatGPT Work, and its speed claim relative to Sol Standard.

    Video from @thsottiaux's post
  31. GoodfireOfficialAI score22

    Alzheimer's Translation Challenge opens registration for spring 2027 competition

    AIThe Alzheimer's Translation Challenge, led by Primamente and AlzData with NVIDIA, Hugging Face, Nebius, Talisman Therapeutics, Ultima Genomics, Cellanome, Prime Intellect, and Boltz, is now open for registration. The competition starts in spring 2027, according to the post.

  32. GoodfireOfficialAI score25

    Goodfire launches a challenge to build AI models on a dataset

    AIParticipants will use a dataset to build AI models, evaluated through a series of evals ranging from general benchmarks to more complex tasks. Top teams will be shortlisted and have their experimental hypotheses tested in Prima Mente's wet lab.

  33. GoodfireOfficialAI score44

    Alzheimer's Translation Challenge Built on 150M-Cell Atlas

    AIThe Alzheimer's Translation Challenge is built on a new atlas of 150M cells, covering neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations with multi-modal readouts. The data will be made available through the AD workbench and Prima Mente's modeling platform.

  34. Boris PowerXAI score46

    OpenAI's GPT-6.1-Sol leads new Arena Alignment Index for agents

    AIThe Arena Alignment Index, built from over 90K real-world agent sessions across 27 models, ranks OpenAI's GPT-6.1-Sol first with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7. GPT-6.1-Sol also posted the lowest observed rates across the index's three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. The index's authors report that newer models consistently outperform their predecessors across all four labs, suggesting broad progress in agent safety.