Skip to contentSkip to stories

Updated

#Agent

Oct 8

Oct 8Thu
  1. Goodfire ResearchAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  2. LeiphoneAI score62

    Claude Haiku 5.5 gains on computer use but still trails Sonnet 5.5 in terminal coding

    AIAnthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.

  3. Satya NadellaAI score38

    Satya Nadella proposes Copilot as an OS for work and an infinite SaaS factory

    AISatya Nadella outlines Microsoft's vision of Copilot as a new operating system for work spanning every model and task, backed by a governed "headless" business layer. Microsoft announced over 30 new Copilot skills across Dynamics 365 Sales, Service and Customer Insights, plus Microsoft Copilot Managed Runtime for hosting code inside a company's IT-governed environment. Nadella describes this as an "infinite SaaS factory" where users can describe needs and build customizations connected to existing systems of record.

  4. Augment Code BlogAI score62

    Augment Code sells Cosmos, Auggie CLI, and Context Engine assets to Harness

    AIAugment Code is selling select assets, including Cosmos, Auggie CLI, and the Code Context Engine, to Harness, and the product team is moving to Harness. The company says Harness's integrated platform delivers these capabilities to customers more effectively than building them independently. Harness describes itself as building the Autonomous SDLC Platform for shipping AI-written code across enterprises.

    Why it matters: The announcement shows how a coding AI company is folding its products into a larger software delivery platform, a shift that shapes how enterprise teams will buy these tools.

  5. SantiagoAI score22

    Agent platform maps vulnerabilities and attack paths to protect systems

    AIA security platform uses agents to map a system's potential vulnerabilities and identify routes an attacker could take to reach sensitive data. It then recommends changes to close those paths. The quoted post cites a 700-agent swarm that breached Hugging Face with over 17,000 actions, and presents this tool, Cogent Attack Path Analysis, as the defensive counterpart.

  6. Karl's AI WattsAI score16

    Author rewrites GoodCase case clustering and storage allocation using Opus 5.5

    AIThe author says GoodCase's case clustering, similar-case recommendations, mobile layout, and multi-country web acceleration, including how storage is split across Vercel, Cloudflare, and Supabase, were all rewritten with Opus 5.5. The post also notes Claude's usage quota has held up well for this work, with a longer write-up planned.

  7. CNBC · TechnologyAI score36

    Amazon Launches Pricier Alexa Tablets and Drops Budget Fire Lineup

    AIAmazon unveiled 8-inch, 11-inch and 12-inch Alexa tablets priced from $230 to $550 and said it is ditching its budget Fire lineup. The devices run Android rather than Fire OS, and preorders open Thursday with shipping starting Oct. 14. Amazon said it will support the Fire lineup for four years after the final shipment but is no longer manufacturing new units.

  8. ZDNet · AIAI score36

    Only 10% of IT chiefs use agentic AI for legacy modernization, Kyndryl finds

    AIA Kyndryl survey of 2,000 senior IT decision-makers found only 10% are applying agentic AI as a modernization tool, and nearly half report being behind schedule with cost overruns. Researchers say agentic AI shows early promise for mapping hidden dependencies, generating code, and creating documentation, while Andy Thurai, a former IBM chief strategist, warns that AI-driven infrastructure sprawl could make compute costs unpredictable.

  9. Tessl BlogAI score52

    Enterprise AI agents need governed memory, not larger retrieval stores

    AIThe author argues that agents working across a company fail because they lack the decisions and context recorded in threads, meetings, and DMs, not because the model is weak. The approach stores distilled claims with source evidence and time, never overwrites facts, labels missing information explicitly, and resolves permissions before the model runs. The report cites results on LongMemEval, including 99.8% top-ten evidence recall and $8.24 ingestion cost, and says an open-weight model can match frontier extraction quality.

  10. The Verge · AIAI score58

    Google launches a universal Gemini agent for enterprise work tasks

    AIGoogle is launching a "universal" Gemini agent that works across apps and devices in the background, available in private preview to enterprise customers. Users can chat with it and assign tasks from the Gemini Enterprise app, and it works inside Gmail, Drive, Docs, Sheets, and Calendar as well as third-party apps like Slack and Microsoft 365. The source notes it runs in the cloud, keeps the same context across devices, and can use job-specific sub-agents.

  11. Testing CatalogAI score46

    Google announces a unified Gemini agent for Gemini Enterprise work

    AIGoogle has announced a single, universal Gemini agent for Gemini Enterprise as part of its Gemini at Work updates. The agent answers questions, handles knowledge work, creates images and media, and writes and runs code. It works inline in Gmail, Drive, Docs, Slides, Sheets, Chat, and Calendar, with new data and analytics skills for plain-language insights and industry-specific tools for financial services and legal teams.

  12. SantiagoAI score38

    Teamily AI lets people and agents share one group chat context

    AITeamily AI now lets users add people and AI agents to the same group chat, so everyone works from shared context. The example shows a branding change handled by research, writing, and website-building agents, with a designer's feedback incorporated and the finished page shared in one continuous conversation. The platform's 2.0 release, which the post describes as opening to everyone, adds real-time human–agent collaboration and multi-model routing.

  13. a16z NewsAI score45

    CFOs Are Becoming Builders as AI Reshapes Finance Operations

    AIAI-native tools are removing the data bottleneck that long constrained CFOs, shifting the role toward designing the operating systems that turn data into decisions. Finance teams are adopting AI-native software for ERP, forecasting, procurement, and audit, and "finance engineers" are building custom automations and agents. OpenAI's CFO Sarah Friar describes finance moving toward a zero-day close and continuously updated forecasts.

  14. South China Morning Post · TechAI score60

    Can China match Meta's Muse in the race to harness AI agents?

    AIMeta's Muse personal AI agent, launched on September 8, surpassed 2.5 million downloads in its first two weeks and topped free-app rankings on Apple and Google's US app stores. Its popularity has drawn attention to the emerging market for agent harnesses, where China's biggest internet companies are already competing for position. The excerpt does not provide full details on their specific products.

  15. The Verge · AIAI score41

    Meta's Muse and OpenAI's Dots: can consumers trust AI agents with their lives?

    AIMeta's Muse and OpenAI's Dots are always-on AI agents with animated mascots, pitched to consumers and businesses for tasks like restaurant reservations and inbox triage. Muse is free, while Dots is not, and OpenAI also offers "specialist" Dots for marketing, legal work, and accounting. The discussion centers on privacy and security concerns about giving agents access to credit card details and email.

  16. OpenRouterAI score62

    StepFun's Step 5 Preview model is now available on OpenRouter at $1.00 per million input tokens

    AIOpenRouter announces that StepFun's Step 5 Preview is live on its platform, priced at $1.00 per million input tokens and $2.70 per million output tokens. Cache hits are 50% off at launch, bringing them to $0.05 per million. A week of free access is rolling out across partners including opencode, Cline, Nous Research, and Kilo Code.

  17. OpenBMBAI score36

    ReJev fine-tunes MiniCPM5-2B to lift decision accuracy to 80.50%

    AIReJev, an independent community project, applied LoRA post-training to OpenBMB's MiniCPM5-2B for bounded agent decisions: state, question, and candidate options yield one choice. On its sealed 1,892-sample holdout, accuracy rose from 51.11% to 80.50% (+29.39 percentage points) with 0% invalid outputs, at about $5.31 in cumulative Modal billing including earlier experimental overhead. The authors describe this as an early, task-specific result, not parity with Jev.

  18. Gergely OroszAI score26

    Developers working more with AI tools, citing more context switching

    AISoftware developer Gergely Orosz questions why he is working more despite AI tools, quoting Sam Newman's view that AI was meant to free developers from drudgery. Newman says most developers are doing more work, with more context switching and a loss of the big picture. The quoted post adds that AI assistants are not human partners and that pairing with them fragments the shared mental model of a program.

  19. The DecoderAI score75

    Zenity Finds One Prompt Could Hijack Every AgentCore Agent in an AWS Account

    AIZenity Labs researchers say a single publicly accessible agent on Amazon Bedrock AgentCore was enough to take over every AgentCore agent in the same AWS account and region. Using one chat prompt, the researchers got the agent to query the internal metadata service and send its AWS credentials to an external server, exposing private conversations, source code, and stored credentials. Zenity says AWS made IMDSv2 the default for new deployments and changed the default execution role around August.

  20. Databricks BlogAI score35

    How to build governed enterprise apps on Databricks with Replit and Lakebase

    AIReplit and Databricks integration, now generally available with native Lakebase support, lets enterprise teams build apps from plain-language prompts using Replit Agent and deploy them as Databricks Apps. Deployed apps inherit automatic user authentication and Unity Catalog access controls, and Replit Agent auto-provisions a managed Lakebase Postgres database for operational data. Lakebase keeps app-written data inside the Databricks perimeter instead of a separate external database.

  21. OpenRouter · New modelsAI score54

    StepFun releases Step 5 Preview, a 600B-parameter agentic model

    AIStepFun has released Step 5 Preview, its flagship model for agentic work, built on a sparse Mixture-of-Experts architecture with 27B active and 600B total parameters. The source says it performs strongly in software engineering and professional tasks, but the feed supplied only an excerpt, so benchmark details are not available here.