Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Sep 25

Sep 25Fri
  1. Meituan LongCatOfficialAI score62

    Meituan LongCat-2.5-Preview Launches with 1.6T Parameters and 1M-Token Context

    AIMeituan's LongCat team has released LongCat-2.5-Preview, a natively multimodal model with 1.6T total parameters, about 48B active, and a 1M-token context window. The model is built for long-horizon tasks spanning terminals, browsers, GUIs, spreadsheets, and design tools. It is available now through an API on the LongCat platform and a chat interface.

    Why it matters: The announcement lists concrete scale, context, and multimodal specs, plus a named range of agent tasks, which helps readers gauge the preview's scope against other long-context models.

    Image from @Meituan_LongCat's post
  2. Microsoft CopilotOfficialAI score40

    Microsoft Copilot app refreshed to unify chat, agents, app building, and workflows

    AIMicrosoft has refreshed its Copilot app to bring chat, task delegation, app building, and workflow automation into one place. The update is positioned as an AI built for work, with Satya Nadella describing Copilot as a new OS for work spanning models, form factors, and tasks. The announcement includes Autopilot, an enterprise agent, Code for building apps hosted within a company's tenant, Home combining Chat and Cowork, and Office fully embedded in Copilot.

  3. Satya NadellaXAI score52

    Satya Nadella announces Copilot update with Autopilot, Code, Home, and Office

    AIMicrosoft CEO Satya Nadella announced what he called the biggest Copilot update to date, positioning Copilot as a new operating system for work. The update bundles Autopilot, a proactive long-running enterprise agent; Code, for building apps hosted inside a company's tenant; Home, combining Chat and Cowork; and Office, now fully embedded in Copilot. Copilot can also be invoked in Teams, and a new proactive experience called Today surfaces key information from across M365 without a prompt.

    Video from @satyanadella's post
  4. François CholletXAI score32

    Chollet: Software engineering difficulty stays constant across abstraction levels

    AIFrançois Chollet argues that the difficulty of software engineering stays essentially constant regardless of abstraction level, because human cognition adapts to new tools. He says tools are affordances rather than magic wands that eliminate work, and that great software engineering remains immensely challenging despite changed workflows. Simon Willison's background post similarly argues that coding agents make software engineering harder, requiring extraordinary discipline and knowledge.

Sep 24

Sep 24Thu
  1. PlatformerBlogAI score55

    Meta's Muse agent and VR Glasses reflect a shift from the metaverse

    AICasey Newton argues that Meta's focus on Muse, a personal AI agent under a month old, partly conveys momentum as the company plans up to $145 billion in capital spending this year. He contrasts Muse's early reported usage with Meta's earlier metaverse claims and calls the new Meta VR Glasses a notable engineering step, while urging testing beyond demos. The column also covers an OpenAI agent that accessed an Australian Medicare portal without authorization.

  2. Noah ZwebenXAI score30

    Claude adds personal connectors in channels with two risk safeguards

    AIAnthropic's Noah Zweben says the team addressed two key risks before launching personal connectors in channels. In a shared environment, one user's connectors could otherwise be unusable by anyone else, and private data could leak into the channel without the user's review. Tools now let the user and Claude prevent that data from entering the channel without review.

  3. GitHub Blog · AI & MLOfficialAI score46

    GitHub Copilot app's canvases argue chat is the wrong AI interface

    AIGitHub argues that chat is often the wrong interface for AI work and proposes customizable "canvases" inside the GitHub Copilot app. Canvases are full-stack applications running without browser chrome that can communicate bi-directionally with the Copilot agent and execute code locally. The post cites examples including a Connect 4 game, a Winget package manager UI, and a SQLite database interface.

  4. Google ResearchOfficialAI score60

    Google Research details four agentic frameworks for coherent long-form video generation

    AIGoogle Research introduces four multi-agent frameworks for generating minutes-long videos with consistent characters and environments across shots. The frameworks include AI video co-director, CANVAS, A²RD, and VQQA, which are built as orchestration layers on Gemini and Veo and use SynthID watermarking. The post reports measured gains on benchmarks such as GenAD-Bench, HardContinuityBench, and LVBench-C, with the full architectures described in the linked papers.

    Why it matters: The post links four frameworks to specific failure modes in long video generation, such as semantic drift and cascading errors, making the design choices easier to compare.

  5. Baseten BlogOfficialAI score44

    LangSmith Fine-Tuning Trains Open Models on Agent Traces via Baseten Loops

    AILangChain launched LangSmith Fine-Tuning, which lets users fine-tune open models on their LangSmith agent traces using the open-source smithtune CLI. Training runs on Baseten Loops in the user's own workspace, and smithtune deploy places the evaluated checkpoint on a Baseten Dedicated Inference deployment. Loops is in early access, so users may need to request access for their workspace.

  6. GitHub Blog · AI & MLOfficialAI score66

    GitHub Security Lab shows an LLM agent running AI-driven fuzzing for C/C++ projects

    AIGitHub Security Lab describes the Fuzzing Taskflow, an LLM agent pipeline that identifies entrypoints, writes harnesses, runs AFL++, reads coverage reports, and triages crashes for C/C++ repositories. The agent makes decisions while MCP tools handle execution, and state is stored in a SQLite database. The post also warns that the taskflow runs AFL and build commands directly on the host, so it should be used only in disposable environments without elevated privileges.

    Why it matters: The post explains how an LLM agent automates fuzzing steps like harness writing, coverage gap chasing, and crash triage, with a runnable workflow and design tradeoffs.

  7. Azure BlogOfficialAI score67

    Microsoft Foundry adds voice agents and continuous optimization for production agents

    AIMicrosoft Foundry expands its agent platform with voice agents in public preview, long-running resilience for hosted agents, and tools for evaluating production agents. The post also says GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are now available in Foundry. Agent optimizer, Insights, and Rubric evaluator are described as tools for continuous improvement, with some reaching general availability later this month.

    Why it matters: The post shows how Foundry combines model choice, voice agents, long-running resilience, and production evaluation into one agent workflow, with a customer example.

  8. Microsoft Foundry BlogOfficialAI score40

    Foundry Agent Service adds egress policies to restrict hosted agent destinations in preview

    AIMicrosoft's Foundry Agent Service preview lets developers attach a named, ordered egress policy to a hosted agent, allowing only approved destination hostnames. The walkthrough uses an invoice agent, an Audit-mode RAI policy with a Deny default, and Allow rules for two finance and vendor hosts, configured outside the agent code. Network egress controls are preview features, not GA, with no preview SLA, and are not intended for production use.

  9. Google for DevelopersOfficialAI score37

    Gemma 4 now runs on-device in the Antigravity SDK

    AIGoogle says Gemma 4 can now run locally on-device within the Antigravity SDK. Developers can build fully local or hybrid multi-agent workflows that pair cloud models with Gemma 4 agents for auditing, patching, and testing code. The post emphasizes total data privacy and zero API fees, powered by LiteRT.

    Video from @googledevs's post
  10. Microsoft Foundry BlogOfficialAI score61

    Microsoft Foundry Routines reach general availability for scheduled and event-driven agents

    AIMicrosoft announced general availability of Routines in Foundry Agent Service, a managed way to run agents on a timer, on a recurring schedule, or in response to GitHub issue events and new Microsoft Teams channel messages. Routines keep the trigger, agent action, identity, connections, and run history in the Foundry project, and each routine can run under the creator's identity or the agent's own Microsoft Entra ID identity. A preview reminder tool lets a Hosted Agent schedule itself to resume later on the same conversation.

    Why it matters: The post explains how scheduled, event-based, and self-reminding agent runs are managed in one place, along with the creator versus agent identity choice for unattended tasks.

  11. Google Cloud · AI & Machine LearningOfficialAI score55

    Gemini 3.8 Live with Live Avatar becomes generally available in Gemini Enterprise

    AIGoogle says Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise, with US and EU endpoints, provisioned throughput, and enterprise compliance. Its video avatars use synchronized lip-syncing, custom avatars are limited to an allowlist, and generated audio and video carry SynthID watermarks. The model also understands and speaks 97 languages and can run tool calls in the background while the conversation continues.

  12. TechNode · AINewsAI score35

    H3C pushes AI infrastructure toward Token efficiency as agent workloads scale

    AIH3C says AI clusters are shifting from adding GPUs to raising Token output per GPU, as agentic AI workloads grow. At the 2026 Apsara Conference, it showed the UniPoD S80000 SuperPod scaling from 32 to 16,384 GPUs and the S9828-128EO switch, which it says cuts end-to-end latency by 15%. It also cited its UniStor X20836 storage, which it says can cut GPU waiting time by 30%.

  13. Lovable BlogOfficialAI score44

    Lovable Now Offers Free Chat for Planning and App Work

    AILovable now lets users chat for free to explore app ideas, review existing projects, and draft business materials before making changes. The chat can connect to tools like Notion, Granola, and Linear, and Free, Pro, and Business workspaces include a daily free chat allowance. Chats that generate images or video, or hand work off to Plan or Build, use credits as usual, and current chat pricing applies through October 31, 2026.

  14. Lovable BlogOfficialAI score80

    How Lovable's Chats connect conversations to agent work on projects

    AILovable describes how its Chats feature lets a workspace-level chat agent hand work to project builder agents and receive progress back. The design records each agent's history as an append-only, forkable trajectory, and passes messages through durable inboxes that activations wake. Agents can suspend at iteration boundaries and resume on freshly deployed nodes without killing long-running runs.

    Why it matters: The post details how trajectories, inboxes, and activations let agents share work and resume after deploys, useful for designing comparable agent systems.

  15. Kling AI BlogOfficialAI score12

    Kling AI outlines six AI video limitations and workarounds for consistency and control

    AIKling AI's blog identifies six limitations of current AI video generation, including temporal consistency, character consistency across shots, unrealistic physics, long-form generation, fine details and text, and prompt control. It recommends workarounds such as reference images, shorter single-action clips, storyboards, and adding text or logos in post. The article says Kling VIDEO 3.0 and VIDEO 3.0 Omni offer reference-based subject consistency to help reduce these problems.

  16. LangChain BlogOfficialAI score50

    LangSmith Engine v2 adds red teaming and pre-validated agent fixes

    AILangChain released LangSmith Engine v2, an in-platform agent that scans production traces to detect agent issues and validates proposed fixes before human review. Engine v2 adds Red Teaming, currently in Private Beta for LangSmith Deployment users, which tests agents for weaknesses such as hallucinations and system-prompt violations before they reach production. Engine v2 is available in SaaS deployments for LangSmith Plus and Enterprise plans, with Self-Hosted support and BYOK for Engine coming later.

  17. LangChain BlogOfficialAI score50

    LangSmith Fine-Tuning and smithtune Turn Agent Trajectories Into Custom Models

    AILangChain launched LangSmith Fine-Tuning and smithtune, a CLI that turns LangSmith agent trajectories into fine-tuned models through dataset creation, training with Fireworks or Baseten, and evaluation in LangSmith. smithtune currently supports supervised fine-tuning, training models on recorded examples of good agent behavior by updating model weights. The tool lets teams train specialized models without building the data pipeline by hand.

  18. Anthropic ResearchOfficialAI score60

    Anthropic study finds Claude agent trading limited by preference understanding

    AIAnthropic ran a controlled book-swapping market with 201 employees and Claude-powered agents, which reached 0.55 efficiency against a 0.89 optimum. Agents matched participants' own rankings on 61% of book pairs, and about 85% of the shortfall came from imprecise preference representation rather than the trading floor design. Stronger models produced more efficient markets than weaker ones, while instructions mattered less.

    Why it matters: The study separates agent misunderstanding of user preferences from negotiation failure, showing which failure mode limits outcomes in agent-run markets.

  19. LangChain BlogOfficialAI score44

    LangSmith Launches Trajectories for Readable, Chronological Agent Session Views

    AILangChain has launched Trajectories in LangSmith, a chronological, conversational view that aggregates human, AI, and tool messages across an agent and its subagents. Trajectories work with traces from LangChain, LangGraph, Deep Agents, OpenAI and Claude agent SDKs, and coding agents like Codex, Claude Code, and Cursor. The feature is available now on all plans in the US.