Gemini 3.5 Flash is now available globally
AIGoogle's Gemini 3.5 Flash is available today globally, according to Oriol Vinyals. The post invites developers to build agents and apps with it and links to Google's blog for more details.
Updated
Updated
Showing low-relevance items too. Hide low-relevance items
AIGoogle's Gemini 3.5 Flash is available today globally, according to Oriol Vinyals. The post invites developers to build agents and apps with it and links to Google's blog for more details.
AIGemini 3.5 Flash produces fully tactile, interactive HTML/SVG hardware simulations in a single shot when given complex agents and tasks. The output includes bump-mapped metals, spring physics, and procedural audio, according to the post.
AIGoogle's Gemini 3.5 Flash reaches 76.2% on Terminal-Bench 2.1, 1656 Elo on GDPval-AA, and 83.6% on MCP Atlas, showing strong performance on multi-step agentic tasks. The model is positioned as an ultra-fast reasoning engine for powering AI agents.

AIComposer 2.5 is now the most-chosen model in Cursor, according to Cursor co-founder Michael Truell. To celebrate, Cursor is giving users 10x usage for the rest of the day.
AIMichael Truell of Cursor says Composer 2.5 is a significant step up from Composer 2. He adds that this is only the start of work with SpaceXAI, with more improvements expected soon. Cursor's announcement describes the model as more intelligent, better at long-running tasks, and more reliable at complex instructions, with doubled included usage for the next week.
AIEugene Yan shares Cloudflare's description of a vulnerability discovery harness that runs eight stages, from reconnaissance to report writing. The pipeline uses about 50 concurrent agents to hunt for bugs, independent agents to try to disprove findings, and a trace step to confirm whether attacker input reaches each bug. Reachable findings feed back into new hunt tasks before a report is written against a predefined schema.

AICognition has released Auto-Triage in Devin Automations, which lets Devin respond to Slack messages, Linear events, GitHub activity, schedules, and webhooks. Devin can investigate with connected observability tools and the codebase, then post a summary, tag an owner, or open a PR. Devin runs in network-sandboxed environments with added protections against prompt injection and data exfiltration, and a limited-time offer gives $200 in credits for a first automation.
Why it matters: The post shows how an agent handles alerts and bug reports from existing team channels, a practical pattern for teams weighing automated incident response.
AIAmazon has rolled out Alexa for Shopping to all customers, combining its Rufus and Alexa+ assistants into a single AI shopping tool. The assistant remembers users' preferences, past purchases, and conversations, and carries that context across phones, laptops, and Echo devices.
AIEugene Yan relays two evaluations of Mythos: UK AISI reports it completed a 32-step network attack, estimated at about 20 expert hours, in 6 of 10 tries and was the first model to solve its end-to-end cyber ranges. XBOW's evaluation describes its performance as token-for-token and unprecedented in precision. The post links both AISI and XBOW blog posts for details.
AISoumith Chintala posted more demos of Interaction Models collaborating live on system design, paper reading, and fact-checking with generative UI. A quoted demo shows the model seeing the user's screen and drawing on it together while building a scalable system architecture.
AIAwni Hannun says Claude Code's new agent view is where he starts and manages much of his work, calling it an exceedingly useful feature. Anthropic's Claude account introduced agent view as a single list of all sessions, available now as a research preview.
AIThinking Machines, founded to advance human-AI collaboration, says its first bet is interactivity built into the model rather than added as scaffolding around a turn-based core. The company argues that how people work with AI matters as much as how intelligent the model is, and that interactivity should scale with intelligence. The post links to a blog detailing these interaction models.
AIMira Murati's post announces interaction models, a new class of model trained from scratch to handle real-time interaction natively rather than adding it onto a turn-based model. The post links to a video, but it provides no benchmarks, parameter counts, or availability details.
AICognition reports that deploying its Devin agent on hardware-in-the-loop and software-in-the-loop workflows cut failure triage time and multiplied test generation at automotive customers. One team reclaimed 2K–4K engineering hours monthly across about 4,000 tickets, while RV Tech rose from 1–2 to 10–15 generated tests per day. Devin also helps convert bottlenecked HIL tests into SIL equivalents to catch failures earlier.
AIBaidu's PaddlePaddle account announced ERNIE 5.1, which it says cuts total parameters to about one-third and activated parameters to about one-half, using roughly 6% of the pretraining cost of models at similar scale. The post reports benchmark results including 99.6 on AIME26 with tools, surpassing DeepSeek-V4-Pro on τ3-bench and SpreadsheetBench-Verified, and ranking #4 globally on Arena Search. ERNIE 5.1 is available through the ERNIE website and Baidu AI Studio Model Playground.
AIMETR's external review of the "Risks from automated R&D" section of Anthropic's February 2026 Risk Report concludes the report does not adequately support its claim that catastrophic risk from Claude Opus 4.6 or a less capable model automating R&D is very low.
AIEugene Yan outlines five principles for working effectively with AI models: treating context as infrastructure, taste as configuration, verification as the basis for autonomy, scaling through delegation, and closing the loop. The post is a short list of themes linked to a longer essay, and no further detail is given in the post itself.
AIOpenAI announced that users can sign in to OpenClaw with their ChatGPT account and use their subscription there. Sam Altman's background post says the same, confirming the ChatGPT login and subscription access for OpenClaw.
AIOpenAI released Auto-review in Codex, which replaces user approval at the sandbox boundary with a separate agent that approves or denies boundary-crossing actions. In internal deployment, Codex sessions stopped for human approval about 200x less often than in manual mode, and Auto-review approved around 99% of escalated actions. The post also states that Auto-review is not a guarantee of security and cannot protect against model scheming.
Why it matters: The post explains how Auto-review replaces human approval at the sandbox boundary, with internal deployment figures and stated limits that help readers judge the tradeoff for coding agents.
AIAndrej Karpathy argues that LLMs enable new functionality, not just faster versions of existing software, citing an image-to-image app, markdown skill files replacing bash install scripts, and LLM knowledge bases over unstructured data.
AIAndrej Karpathy describes a December 2025 shift in which coding agents began producing larger, more reliable chunks of work, changing programming toward orchestrating agents. He argues that models automate what can be verified and that their capability is jagged, depending on verifiability and what labs emphasize in training, so users need to stay in the loop. He also says hiring, founder opportunities, and agent-native infrastructure should adapt to this shift.
AICognition has opened its Asia-Pacific headquarters in Singapore to expand its autonomous software engineering platform, Devin, across the region. The company says OCBC saw up to 30% improvement in code and test case generation, and its system integration test first-pass rate rose from below 50% to over 80% after deployment. Cognition is building its Singapore team across engineering, go-to-market, and partnerships, with Richard Spence leading APAC.
AINVIDIA's developer account announced a livestream titled "Scaling NemoClaw: Roadmap, OpenClaw Collaboration, and Real-World Integration" as part of Nemotron Labs. The post provides no further details on the roadmap, collaboration terms, or integration specifics.
AIMercedes-Benz is deploying Cognition's Devin and Windsurf across its global engineering teams, from the United States to Europe and Asia. In a four-week pilot, Devin analyzed over 200,000 lines of COBOL code and cut modernization time from an estimated eight months to eight days. The company is now rolling out the full suite, with Windsurf for development, Devin as an autonomous cloud agent, and Devin for Terminal for the most complex tasks.
AIAndy Jassy announced that OpenAI's models will become available directly to customers on Amazon Bedrock in the coming weeks. The availability will be alongside the upcoming Stateful Runtime Environment, giving builders more model choice. More details are expected at the AWS event in San Francisco tomorrow.
AIXiaomi's MiMo-V2.5 is a native omnimodal model that understands text, image, video, and audio within one architecture. It is a sparse MoE with 310B total and 15B activated parameters, and supports up to 1M tokens of context. The repository also notes a config.json and tokenizer_config.json update that users who downloaded before commit 4da2748 should re-pull.
Why it matters: The repository documents a 310B-parameter omnimodal MoE with a hybrid attention design, useful for comparing long-context efficiency against other open multimodal models.
AIMistral AI has released mistralai/Mistral-Medium-3.5-128B-EAGLE, an EAGLE draft model for speculative decoding with the 128B dense Mistral Medium 3.5. The companion model, which the source says replaces Mistral Medium 3.1 and Magistral in Le Chat and Devstral 2 in Vibe, has a 256k context window, handles text and image input with text output, and is served with vLLM or SGLang using three speculative tokens. The model is released under a Modified MIT License that allows commercial use with exceptions for companies with large revenue.
AISoumith Chintala comments on a Reddit report that Anthropic banned organizations without warning, suggesting Anthropic may need to scale Account Support using Claude or human account managers. He also argues that enterprises may increasingly adopt multiple AI providers with open harnesses, facing cloud-era vendor problems that would likely affect all AI providers.
AIXiaomi released and open-sourced MiMo-V2.5-Pro, a 1.02T-parameter Mixture-of-Experts model with 42B active parameters and a 1M-token context window. The company reports gains in agentic tasks, software engineering, and long-horizon work, including a Rust SysY compiler task finished in 4.3 hours across 672 tool calls. Weights and tokenizer are on Hugging Face, and API pricing is unchanged.
Why it matters: The release pairs a 1.02T-parameter open-weight model with long-horizon agent results and token-efficiency claims, useful for judging its fit in coding and agent workflows.
AICognition has released Devin CLI, a local coding agent that runs in the shell with access to the codebase, tools, and environment. Users can choose among frontier models including Opus 4.7, GPT-5.5, and SWE-1.6, and hand sessions off to cloud agents with their own computer, which keep working after the laptop is closed.
AIMckay Wrigley says his coding split moved from 80/20 Claude/GPT to 80/20 GPT/Claude within three months, and he trusts GPT 5.5 for engineering. He still finds Claude better for non-coding agent work, calls Opus 4.7 underwhelming, and attributes Anthropic's issues to compute constraints.
AIAhmad Al-Dahle argues that the most interesting part of DeepSeek-V4 is its bet on efficient ultra-long context rather than its benchmarks. He says this is the precondition for test-time scaling and long-horizon agents, and cites 27% of V3's FLOPs at 1M tokens. The quoted DeepSeek post announces DeepSeek-V4-Pro (1.6T total, 49B active) and DeepSeek-V4-Flash (284B total, 13B active), both open-sourced with 1M context and API access.
Why it matters: The post argues that efficient 1M-token context, not benchmark scores, is the key bet behind DeepSeek-V4's design for test-time scaling and long-horizon agents.
AIFactory has released an Automated QA skill that drives an app as a real user would, filling forms, typing into terminals, and calling endpoints, then posts a structured report with screenshots, terminal snapshots, and API traces as a single updating comment on each pull request. Teams can run it on every push or make it an optional CI check triggered by a PR label, comment command, or manual dispatch, and developers can run /qa locally in any Droid session. Automated QA is available today in all Factory plans.
AICognition argues that enterprises building cloud agents face three problems: shared container kernels, the inability to persist agent state across async gaps, and the scale of orchestration, governance, and integrations. The post says VM-level isolation with hypervisor-level snapshots was needed for Devin, and that organizations must also rebuild engineering processes around agent execution.
AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.
Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.
AIXiaomi released MiMo-V2.5, a 310B-parameter sparse MoE model with 15B active parameters that adds native visual and audio understanding. The model supports up to 1 million tokens of context, and its weights, tokenizer, and model card are available on Hugging Face. Xiaomi says it surpasses MiMo-V2-Pro on agentic performance and reports a Claw-Eval score of 62.3 on the general subset.
Why it matters: The release pairs native visual and audio understanding with a 1M-token context window and open weights, a combination worth checking against your own multimodal workflows.
AICursor's Michael Truell says the company is partnering with the SpaceX team to scale up Composer, calling it a meaningful step toward building the best place to code with AI. The quoted SpaceX post says Cursor gives SpaceX the right to acquire Cursor later this year for $60 billion, or pay $10 billion for the work together. It also cites SpaceX's Colossus training supercomputer, described as a million H100-equivalent system, as a source of training capacity.
AINVIDIA's OpenShell v0.0.34 release lets users update sandbox policy without restarting the runtime. The update also adds install-vm, which installs the gateway and VM driver with new --driver-dir support, and sandbox get, which shows the active runtime policy. Supervisor seccomp improvements and HTTP normalization are included as well.
AINVIDIA released OpenShell v0.0.33, which adds seccomp and process-limit hardening, inference routing, and a standalone libkrun compute driver. The libkrun driver provides a lightweight VM backend as a second compute path for agents. The release also includes bug fixes, a docs refresh, and improved test stability.
AINVIDIA released OpenShell v0.0.32, which ships standalone openshell-gateway binaries so the gateway can be deployed without building from source. The release also adds system CA certificates for upstream TLS and reorganizes RFC documentation by RFC.