Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. OpenAI · YouTubeOfficialAI score24

    How Oracle Uses ChatGPT Work to Transform Recruitment Planning

    AIOracle built a talent market intelligence tool with ChatGPT Work to transform hiring preparation, according to Jan Ackerman. Starting from a job description, the tool researches comparable roles, benchmarks compensation, and assesses talent pools across locations to give hiring managers consistent data and insights.

  2. OpenAI · YouTubeOfficialAI score67

    OpenAI launches GPT-6 Intelligent UI for interactive ChatGPT answers

    AIOpenAI's GPT-6 in ChatGPT adds Intelligent UI, which lets ChatGPT answer with interactive interfaces and build quick tools for a task. The feature is rolling out to Plus, Pro, Business, and Enterprise first, with Free and Go tiers following, and Enterprise access depends on workplace admin settings. The update covers only the Chat experience, and the models powering Work and Codex are not changing.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”

  3. ClineOfficialAI score46

    Cline makes Step 5 Preview free, citing strong DeepSWE coding scores

    AICline says Step 5 Preview is now free in its coding tool and scores ahead of Kimi K3 and GLM-5.3 on DeepSWE. The company describes it as one of the strongest open-weights coding models available. StepFun's background announcement describes Step 5 Preview as a 600B total / 27B active MoE model with 1M context and vision, and says open weights arrive on Oct 15.

    Image from @cline's post
  4. Sierra BlogOfficialAI score62

    Sierra launches fleming-1 to detect AI agents calling by phone

    AISierra has launched fleming-1, a model that analyzes caller speech in real time and scores audio for signs it was generated by AI. It flags likely AI callers while keeping real people unflagged by default, and companies decide how to handle those calls. The model works with any voice agent built on Sierra, and Sierra also announced Personal Agent Protocol, an open standard for authorized agent-to-business interactions.

    Why it matters: The post explains why companies need to know when a caller is an AI agent, which frames the detection model as a business decision rather than an automatic block.

  5. 🚨 AI News | TestingCatalogXAI score49

    Odyssey launches Odyssey-3 world model with public research preview

    AIOdyssey has launched Odyssey-3, its most powerful foundation world model, with a public research preview. Odyssey-3 Pro scored 66.1 on Physics-IQ Verified video-to-video with best-of-8 sampling, the highest reported result. The model generates environments from prompts and predicts changes in real time as users move through scenes.

    Image from @testingcatalog's post
  6. elvisXAI score42

    Odyssey-3 Pro tops Physics-IQ Verified and shows robotic error recovery.

    AIOdyssey released Odyssey-3 Pro, which sets a new top score on Physics-IQ Verified, a benchmark where models continue videos of real physics experiments. In robotics, a robot arm with tens of hours of demonstrations recovered from a missed grasp, a behavior absent from those demos.

    Video from @omarsar0's post
  7. Arena.aiOfficialAI score37

    Arena raises $200M Series B at $3.1B valuation, launches Alignment Index

    AIArena announced a $200 million Series B at a $3.1 billion valuation, alongside a new Alignment Index that measures whether AI agents behave safely, truthfully, and within the bounds of user requests. The company has surpassed $100 million in annualized revenue, facilitated 350 million sessions and 62 million votes, and led by Felicis and PXD from the seed and Series A stages. Arena positions the index as a way to assess trustworthiness as AI systems increasingly take real actions.

  8. Will HunterXAI score52

    Cognition built Devin as a marketing ops manager using tested code and approval gates

    AICognition engineered its marketing operations around Devin by building each workflow in code with tests, so Devin can run and debug it. Devin connects to nine systems, including Salesforce, HubSpot, Meta Ads, and LinkedIn Ads, and each change shows the exact edit, waits for a confirmation phrase, and reads the result back before reporting it. Cognition says it is hiring marketers and GTM engineers to extend the system.

  9. SiliconANGLE · AINewsAI score30

    Liquid AI Builds On-Device Personal AI Around Device-Level Context

    AILiquid AI is building personal AI that runs on devices such as phones, wearables, PCs, and cars, using its Liquid Context layer, which is optimized for Snapdragon processors, to sit between models, agents, and hardware. The company's agent harness uses its own models to decide which user context to retain and how to compress it within fixed compute limits. Liquid AI is also collaborating with Mercedes-Benz Group AG to bring on-device AI to its cars and plans observability and continuous improvement loops for self-improving agents.

  10. The Verge · AINewsAI score62

    Anthropic updates Claude usage policy to ban abusive treatment and expand misuse rules

    AIAnthropic is revising its usage policy for the first time in over a year, adding bans on sustained abusive or cruel behavior toward Claude and on deceptive election and propaganda campaigns. The update also expands weapons restrictions, tightens surveillance bans, and requires a qualified operator able to stop equipment when Claude controls autonomous physical hardware. Terminating conversations remains the primary enforcement mechanism, and the company did not say whether user bans would follow.

  11. SantiagoXAI score46

    Odyssey 3 Pro world model tops Physics-IQ and goes live

    AIOdyssey 3 Pro, a world model, is now live as a research preview and ranks first on the Physics-IQ Verified video-to-video benchmark. The post says it can learn from visual observations and map that knowledge to physical controls for robots, cars, video games, and drones. Odyssey-3, the model launched alongside it, is described as free to try.

    Image from @svpino's post
  12. Tessl BlogOfficialAI score44

    Continuous AI Brings Agentic Automation to Repository Workflows

    AITessl's blog post argues that repository automation needs Continuous AI, a third pillar alongside CI and CD for scheduled, auditable AI workflows that improve repositories over time. The article describes GitHub Agentic Workflows, which harden agentic workflow specifications into GitHub Actions that can run coding agents such as Claude Code, Copilot CLI, Gemini CLI, or Codex-style agents. It emphasizes read-only agent steps, restricted outputs, and human review of pull requests.

  13. Meta NewsroomOfficialAI score22

    Meta Debunks Three Common Myths About Its Data Centers

    AIMeta says its closed-loop liquid cooling recirculates water in a sealed system, so its data centers use less water annually than an average US golf course. The company also says it pays for the new generation and transmission its facilities require, including in Louisiana under its Entergy agreement, and that data centers create construction and operations jobs.

  14. OdysseyOfficialAI score31

    Odyssey-3 world model debuts for physical AI and training environments

    AIOdyssey has released Odyssey-3, which it describes as a major leap toward world models that power physical AI, generate training environments, and enable new human experiences. The post invites readers to try Odyssey-3 at the company's website but gives no specific benchmarks, parameter counts, or pricing.

  15. OdysseyOfficialAI score34

    Odyssey-3 world knowledge can be applied to physical AI systems

    AIOdyssey says its Odyssey-3 model's learned world knowledge can be adapted by physical AI developers to control robots, power humanoids, drive cars, and fly drones. The post describes this as a capability for autonomous machines generally, without providing benchmarks, specifications, or availability details.

    Video from @odysseyml's post
  16. OdysseyOfficialAI score38

    Odyssey-3 is a foundation world model for physical AI and agents

    AIOdyssey announced Odyssey-3, a foundation world model it says enables applications in physical AI, human experiences, and training intelligences. The company highlights agents learning from experience inside Odyssey-3 while working toward objectives.

    Video from @odysseyml's post
  17. OdysseyOfficialAI score22

    Odyssey-3 Pro sets new Physics-IQ video-to-video benchmark record

    AIOdyssey-3 Pro achieved a score of 66.1 on Physics-IQ Verified's video-to-video benchmark, the highest reported score so far. Physics-IQ evaluates physical behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.

    Image from @odysseyml's post
  18. Aravind SrinivasXAI score22

    Perplexity Decider ranks first on DecisionBench at lowest cost

    AIPerplexity's Decider V1.1 ranked first on DecisionBench while also having the lowest cost, according to a post highlighting the result. The benchmark results cited include 949 shared text cases, 93.9% accuracy, a 534 ms median latency, and $0.016 per 1k decisions.

  19. LangChainOfficialAI score34

    LangChain's Restock agent buys office supplies through Slack with approval

    AILangChain has built Restock, an office supply agent that works inside Slack and can find real products, prepare purchases, and pay for them. A person approves each order, which is reviewed in Slack and approved through Stripe's Link agent wallet, built on MPP and Managed Deep Agents.

    Video from @LangChain's post
  20. GoodfireOfficialAI score21

    Goodfire's probes run during inference with no added latency

    AIGoodfire reports that running its probes during model inference maintains the same throughput with no added latency. The company attributes this to infrastructure engineering, including kernel-level optimizations and a custom inference server.

  21. Latent SpaceBlogAI score59

    Periodic Labs argues AI scientists need physical experiments, not just more data

    AIPeriodic Labs' Liam Fedus and Ekin Dogus Cubuk explain why scientific discovery differs from math and coding, and why experiments remain the ground truth. They describe reinforcement learning grounded in physical experiments, AI-driven materials characterization, and the view that failed experiments can be valuable training data. The transcript was truncated before the discussion of giving lab instruments "140 IQ" was completed.

  22. SunoOfficialAI score22

    Suno launches Albums for bundling songs into full releases

    AISuno announced that Albums are now live, letting users combine songs into a full release, set artwork, arrange the tracklist, and publish when ready. Existing playlists can be converted into Albums without rebuilding them from scratch.

    Video from @suno's post
  23. Google GemmaOfficialAI score27

    EmbeddingGemma 2 developer guide released by Google

    AIGoogle Gemma has published a developer guide for EmbeddingGemma 2, with code snippets to help developers start searching beyond text. The post directs readers to the full guide on the Google Developers Blog.

  24. Google GemmaOfficialAI score44

    Google publishes a developer guide for EmbeddingGemma 2 multimodal embeddings

    AIGoogle Gemma announces a developer guide showing how to embed text, code, images, video, audio, and interleaved inputs with EmbeddingGemma 2 using the sentence-transformers library. The guide outlines a four-step workflow: loading the model, embedding text and code with task prompts, embedding multimodal inputs, and optionally truncating dimensions with Matryoshka.

    Image from @googlegemma's post
  25. AWS Machine Learning BlogOfficialAI score27

    Share SageMaker HyperPod GPU clusters across teams with isolation and fair scheduling

    AIAWS published a reference architecture for running multiple teams on one Amazon SageMaker HyperPod EKS cluster, with each team isolated in its own Kubernetes namespace. The design combines AWS IAM Identity Center for authentication, per-team SageMaker AI domains, HyperPod Task Governance for fair resource allocation, and namespace-level cost allocation for per-team spend visibility.

  26. MarkTechPostNewsAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.