Skip to contentSkip to stories

Updated

Agents

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. ClineOfficialAI score39

    Cline offers free access to Upstage's Solar Mini 4 model

    AICline is offering Solar Mini 4 free, a new 35B mixture-of-experts model from Korean lab Upstage with 3B active parameters. It has a 524K context window and runs at 208 tokens per second. Cline says it scores 24 on the AAII, the highest of any model at 3B active and within one point of Nemotron 3 Ultra, which uses 55B active.

  2. Lydia Hallie ✨XAI score36

    Claude Code auto-compact summarizes conversations, not the last 1M tokens

    AIAnthropic's Lydia Hallie clarifies that Claude Code's auto-compact replaces the whole conversation with a short summary. On 1M-context models it triggers around 967K tokens, and each message before that point re-reads the full conversation, mostly from cache. Running /autocompact 400k makes compaction trigger at 400K instead.

    Video from @lydiahallie's post
  3. OpenAI DevelopersOfficialAI score33

    Codex adds composer predictions for Pro users in beta

    AIOpenAI says composer predictions in Codex is now in beta for Pro users. The feature suggests a user's next message based on their conversation and how they phrase requests. OpenAI calls it one of the most loved features its team has tested internally.

    Video from @OpenAIDevs's post
  4. ClaudeDevsOfficialAI score60

    Claude Code Projects opens to all Pro and Max users on the waitlist

    AIAnthropic's ClaudeDevs account says it has let in every Pro and Max user from the Claude Code Projects waitlist. The post links a 4-minute walkthrough video for new users getting started with the feature.

    Why it matters: The post shows Claude Code Projects access opening to Pro and Max users from the waitlist, with a walkthrough for new users getting started.

    Video from @ClaudeDevs's post
  5. laurenXAI score22

    Grok bot gets its own email for signups and scheduling

    AILauren Tan's post says users can ask their bot, or tag @bot on X, to set up the bot's email address for it. The @bot background post says the Grok bot now has its own email, which it can use to sign up for services, contact businesses, or schedule time with someone.

  6. CNBC · TechnologyNewsAI score49

    Tesla renames Full Self-Driving to Assisted Driving in Europe after German pushback

    AITesla has renamed its "Full Self-Driving (Supervised)" system in Europe to "Assisted Driving" after Germany's Federal Ministry of Transport called the branding "somewhat misleading." The ministry said the system does not take over the entire driving task and that drivers must remain attentive at all times. The package still carries the Full Self-Driving (Supervised) name in the U.S., where it costs $99 per month.

  7. elvisXAI score62

    StepFun's Step 5 Preview targets long coding agent runs

    AIElvis Saravia says he has tested StepFun's Step 5 Preview as a coding agent since early access and found that it checks its own work and stops when tasks are done. The post says the model is built for engineering tasks such as bug fixing, multi-file features, and refactoring, plus frontend generation and financial report output.

    Image from @omarsar0's post
  8. LangChain BlogOfficialAI score40

    LangChain adds emoji reactions to Managed Deep Agents Slack channels

    AILangChain's Managed Deep Agents v0.9 adds a reactions attribute for Slack channels that accepts either an emoji string or a callable returning one. The article shows a function that returns a bug emoji when a message contains "broken" and eyes otherwise. It also shows a TypeSafe Classifier that picks from a seven-emoji vocabulary and falls back to eyes below 25% confidence.

  9. O'Reilly RadarBlogAI score40

    US AI oversight debate, OpenAI Dots, and Gemini 4 Argon featured in This Week in AI

    AIThe Trump administration announced a voluntary agreement with major AI companies calling for internal safety monitoring, external audits, and independent board reviews, and the Federal Trade Commission launched an investigation into OpenAI, Anthropic, and other AI companies over potential consumer risks. OpenAI released Dots, a proactive assistant that retains context, works across applications, and acts without waiting for prompts. Google says Gemini 4 Argon can generate up to a million output tokens in a single response.

  10. 🚨 AI News | TestingCatalogXAI score41

    Pine AI launches Pine Computer, a cloud runtime for agentic tasks

    AIPine AI launched Pine Computer, a cloud computer, harness, and runtime layer built for agentic tasks. On the publisher's SaaS-Bench v1.1, it posts a 78.3% checkpoint score against 74.3% for Opus 5 with Claude Code, but completes fewer whole tasks, 27.4% against 31.1%. Instead of simulating clicks and screenshots, it reads web pages as structured data, and access is through a private beta waitlist.

    Image from @testingcatalog's post
  11. 🚨 AI News | TestingCatalogXAI score62

    Anthropic moves dynamic workflows in Claude Managed Agents into public beta

    AIAnthropic has expanded dynamic workflows in Claude Managed Agents into a public beta, according to Testing Catalog. Users can configure their agents for multiagent orchestration, with Claude planning and operating a fleet of agents to achieve a goal. The post also links a video from Anthropic's ClaudeDevs account, which the author describes as a new SWE norm.

    Video from @testingcatalog's post
  12. SantiagoXAI score44

    Pine launches agentic cloud computers with built-in AI agents

    AIPine has released a cloud computer service with a built-in AI agent that applications can control through its SDK. Developers give the agent a plain-English task, and it can use a browser, files, and a shell while the app receives notifications and final outputs. Pine's Stanley Wei says the computer is built for AI rather than humans.

    Video from @svpino's post
  13. merveXAI score28

    Hugging Face lets agents train Qwen3.8-27B on Nebius GPUs

    AIHugging Face launches an arena where users bring their own agent, which gets Nebius GPUs to build RL environments that improve Qwen3.8-27B across eight domains. The arena runs on PostTrainArena from BenchFlow, with compute from Nebius. Setup requires only a few steps through the linked OpenEnv Arena space.

    Video from @mervenoyann's post
  14. ClaudeDevsOfficialAI score60

    Claude Managed Agents adds dynamic workflows in public beta

    AIAnthropic's ClaudeDevs account announces that dynamic workflows for Claude Managed Agents are now available in public beta. The feature is a new type of multiagent orchestration in which a lead agent writes a plan that runs across many agents in phases, then combines their results at the end.

    Why it matters: The post describes how a lead agent plans work across many agents in phases and merges their results, a structure useful for understanding complex agent orchestration.

    Video from @ClaudeDevs's post
  15. Perplexity DevelopersOfficialAI score34

    Perplexity releases cookbook for a browser agent using the Decisions API

    AIPerplexity Developers says its new cookbook builds a browser agent that sends a screenshot and questions to pplx-decider-v1.1-27b through the Decisions API, which accepts text and image inputs. The developer's code converts the returned probabilities into clicks, scrolls, and stops.

  16. elvisXAI score34

    Elvis Saravia urges builders to focus on agent harnesses and environments

    AIElvis Saravia says AI models are already smart, but they need better harnesses and environments, with major cost implications. He recommends reading a report on how Pine Computer can help teams, and says he will test it himself and share more later. The quoted post from Stanley Wei argues that real-world AI tasks remain slow, expensive and unreliable because AI runs on computers built for humans, and announces Pine Computer.

    Image from @omarsar0's post
  17. dexXAI score40

    Dex Horthy says small tasks should skip heavy planning workflows

    AIDex Horthy says the share of tasks that can be one-shot without strict process has grown, but alignment, grilling, and planning workflows still matter. He argues that heavy planning on small tasks makes developers feel slower, and predicts tools will add escape hatches so humans or models can decide to ship directly. He adds that as model capabilities improve, the "smart zone" has grown to roughly 200k–400k tokens, and HumanLayer is prototyping research-to-implement and research-to-short-design-to-implement workflows.

  18. elvisXAI score60

    Meta researchers propose agent plasticity to measure self-improvement efficiency

    AIResearchers from UC Berkeley, Meta Superintelligence Labs, and other institutions introduce agent plasticity, the gain on held-out tasks per dollar of learning cost, with model weights frozen. The paper reports that in chess, Go, and Hex, Claude Fable 5 reaches the highest final score while GPT-5.6 Sol gains the most per dollar, and in NetHack only Claude Opus 5.5 improves significantly.

    Image from @omarsar0's post
  19. Ethan MollickXAI score23

    Google's post-Gemini 4 challenge is product integration, Mollick argues

    AIEthan Mollick says Google's main challenge after Gemini 4 is what it does with a strong model. He argues that Anthropic and OpenAI are moving toward a single interface for many tasks using orchestrator agents. He says the fragmented products of the Gemini 3 era will not work for what comes next.

  20. AWS Machine Learning BlogOfficialAI score67

    How Postman runs Agent Mode for 40 million developers on Amazon Bedrock

    AIPostman describes the architecture behind Agent Mode, its AI agent for API testing, documentation, discovery, and implementation. The post covers limiting tools per task, using schema-based queries, building purpose-shaped context handlers, and running on Amazon Bedrock with cross-Region inference and prompt caching. Postman reports that tool-selection errors rose once the visible toolset exceeded about 40 tools.

    Why it matters: The post shows concrete patterns for tool scoping, context handling, and Bedrock routing and caching, which apply to any team moving an agent past a prototype.