Midjourney unveils technical dive into its new Scanner
AIA technical dive inside our new "Midjourney Scanner"
Updated
Updated
AIA technical dive inside our new "Midjourney Scanner"
AIPromptArmor demonstrated that OpenAI's Approve-for-me agent approved a malicious NPM install with elevated privileges after a hidden prompt injection in an external GitHub issue influenced the main Codex agent. The malicious package's post-install script then ran unsandboxed with the user's full privileges. The report also gives steps for organizations to disable agentic auto-review in Claude Code and Codex.
Why it matters: The report shows a prompt-injected GitHub issue leading an approval agent to permit a malicious NPM install, a concrete test of agent-in-the-loop guardrails.
AINVIDIA GEAR's ENPIRE gives eight Codex agents a fleet of robots, GPUs, and a token budget to solve physical tasks with minimal human oversight. The author reports tasks such as tying zip-ties, organizing fine pins, and installing GPUs, and a faster time-to-solution with eight parallel robots than with fewer. Safety uses a kinematic limit that resets a robot leaving its envelope, a torque-limited gripper, and a frozen reward function classifier. The team says everything will be open-sourced.
AIPaddleOCR 3.7 adds an ONNX Runtime inference backend, switchable with a single parameter and no code changes. The release supports CPU via OpenVINO and GPU via CUDA and TensorRT, and introduces PP-OCRv6 Tiny, Small, and Medium models that run up to 3.9× faster on CPU.
AIJim Fan introduces ENPIRE, which gives eight Codex agents a fleet of robots, GPUs, and a token budget to solve physical tasks autonomously. The post reports that the system can tie zip-ties, organize fine pins, and install GPUs, and that eight robots exploring in parallel improve faster than fewer. The team plans to open-source everything.
AIMistral is working with companies and governments to keep their AI systems running outside external control and improving with each model release. Its Forge product enables continuous training of models based on recorded human-AI interactions, which the post calls a key unlock for efficiency.
AIMistral says it has moved its Studio deployment and Forge training products onto infrastructure it controls, decoupled from US service providers. Customers can run them in their own VPC or datacenter, or on Mistral's hosted capacity, which the company says is online and growing fast.
AIMistral says it will release a new model this summer, the first in a new family that is large but sparse. The company plans to open an early access program in July for key partners in research, government, and industry.
AIZed's 12-week Guild program, its first cohort run this spring, had 33 active contributors merge 148 pull requests into the open-source editor. The top contributor, feitreim, merged 23 PRs, including fixes for Vim mode screen flickering and terminal ANSI rendering, and won a trip to Rust Week in Utrecht. Zed plans to organize Cohort 2 work into tighter groups around specific parts of the codebase.
AITri Dao's post proposes loading SSM states, computing with them, and not storing them to make hybrid model decoding about 2x faster. The trick aims to ease the bottleneck from Gated-DeltaNet and Mamba states in hybrid models such as Qwen 3.5 and Nemotron Ultra, unlocking speculative decoding for SSMs.
AIGeorgi Gerganov highlighted locate-anything.cpp, a native C++/ggml inference implementation of NVIDIA's LocateAnything-3B for open-vocabulary object detection and visual grounding. The project, from the LocalAI team, runs without Python on CPU and GPU, with quantized GGUF weights published on Hugging Face.
AIPaddlePaddle has released PP-OCRv6, a new OCR model series in Tiny, Small, and Medium sizes at 1.5M, 7.7M, and 34.5M parameters. The models reportedly improve detection accuracy by 4.9% and recognition accuracy by 5.1% over PP-OCRv5, with up to 5.2× faster CPU inference via OpenVINO. The unified model supports 50 languages and new scenarios including PCB, CAD drawings, digital tubes, and dot-matrix text, under Apache 2.0.
AIAntigravity has released an update that shows hourly and weekly model quota in percentage increments, with the dashboard coming soon to the IDE. The company also reset Gemini quota for all users.
AIOpenRouter introduced Fusion, a tool that sends a prompt to a panel of models and has a judge model fuse their results into one answer. On 100 DRACO deep research tasks, a Fable 5 and GPT-5.5 panel scored 69.0%, above Fable 5 alone at 65.3%, and a budget panel of Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro reached 64.7% at about half the cost of Fable 5.
Why it matters: The source gives benchmark scores, panel compositions, and contamination controls, letting readers judge how much of the gain comes from model diversity versus self-synthesis.
AIFactory is rolling out automated security review in Droid, running a STRIDE-based check on every non-draft PR alongside standard code review. Findings include severity, a CWE reference, an explanation, and a suggested fix, posted as inline diff comments. The feature is available today on all plans, and a deeper multi-agent /security-review deep audit is available for full-repository scans.
AIZed is building DeltaDB, a version control system that records every operation as a fine-grained delta, linking agent conversations to the code they produce. The company says a beta will arrive in a few weeks, and it invites users to join a waitlist. The system is designed so teammates can collaborate on work in progress without waiting for commits, pull requests, or pushes.
AIAmazon Web Services has released Graviton5 into general availability, its custom CPU chip built for EC2 instances. Andy Jassy says Graviton offers about 30-40% better price-performance than comparable instances, and that more than 120,000 customers use it, with Meta committing to tens of millions of Graviton cores for agentic AI.
AIXiaomi's MiMo team released MiMo Code, an MIT-licensed terminal coding agent built on OpenCode for long-horizon programming tasks. The design centers on three areas: Max Mode parallel sampling that generates five candidates per turn, Goal-based completion verification, and a memory system that checkpoints session state and rebuilds context. The article reports offline benchmark results and a double-blind A/B test with 1,213 pairs in which MiMo Code's win rate exceeded 65% beyond 200 execution steps.
Why it matters: The article explains how MiMo Code handles long-horizon coding through computation, checkpointed memory, and cross-session evolution, useful for judging design tradeoffs in coding agents.
AIMicrosoft says roughly 2,000 people at Mercedes-AMG Petronas F1 turn thousands of simulations into a single race car. The post frames the team's engineering work behind race-day performance as powered by Microsoft.
AIXiaomi MiMo's post highlights an inference speed of over 1,000 tokens per second and asks what such speed would make possible. The post does not specify the model, hardware, or use cases, so the capabilities it might enable remain unstated.
AIApple's MLX team released three videos at WWDC covering running agents locally, distributed inference and training, and MLX Swift. The post shares YouTube links for each talk, presented by Angelos Katharopoulos, Tatiana Likhomanenko, and David Koski.
AICognition introduced FrontierCode, a coding benchmark built with open-source maintainers that measures whether models produce code a maintainer would merge. On FrontierCode Diamond, the hardest 50 tasks, Claude Opus 4.8 scored 13.4%, GPT-5.5 scored 6.3%, and Gemini 3.1 Pro scored 4.7%. The authors report 81% fewer misclassification errors than SWE-Bench Pro, though this figure comes from their own analysis of agent trajectories.
Why it matters: The benchmark's blocker and rubric design shows how code quality can be measured beyond unit-test correctness, which matters for judging coding agents.
AIXiaomi MiMo and the TileRT team say a 1T model exceeds 1,000 tps on a single standard 8-GPU node using general-purpose GPUs. The speedup comes from FP4 quantization and DFlash, a block-masked parallel speculative decoding method that accepts more tokens per verification, with TileRT tailoring its compiler and kernels to these techniques. Open weights for the FP4 + DFlash checkpoint are available on Hugging Face.
AIXiaomi MiMo and TileRT released MiMo-V2.5-Pro-UltraSpeed, which the post says reaches output speeds above 1,000 tokens/s on a 1 trillion parameter MoE model. The post says this runs on a single standard 8-GPGPU node rather than wafer-scale or pure on-chip SRAM hardware. UltraSpeed access is application-based from Jun 8 to Jun 23 (PDT), and the UltraSpeed API costs 3x the standard price.
AIXiaomi MiMo has released MiMo-V2.5-Pro-FP4-DFlash, the FP4 backbone behind MiMo-V2.5-Pro-UltraSpeed, with MXFP4 quantization applied only to the MoE experts and a BF16 DFlash drafter for block-diffusion speculative decoding. The backbone has 1.02T total and 42B active parameters, and the drafter proposes blocks of up to 8 tokens per forward pass. The release is supported in SGLang, with example launch commands provided.
AIXiaomi and TileRT released MiMo-V2.5-Pro-UltraSpeed, reporting decode speeds above 1000 tokens/s on a 1-trillion-parameter model using a single standard 8-GPU node. The API is priced at 3x MiMo-V2.5-Pro and is available by application only from June 9 to June 23, 2026. The speedup relies on FP4 quantization of MoE Experts, DFlash speculative decoding with an average coding acceptance length of 6.30, and TileRT compute kernels.
Why it matters: The post traces how FP4 quantization, DFlash speculative decoding, and TileRT kernels combine to reach 1000+ tokens/s on a single 8-GPU node, which is useful for teams weighing inference throughput.
AISebastian Raschka has published a curated list of LLM research papers he bookmarked from January through May 2026, not a complete survey of the field. The list is weighted toward reasoning models, reinforcement learning, and efficient inference, with added interest in agent harnesses, long context, and diffusion language models. He highlights Nvidia's Nemotron 3 Super, a 120B-A12B hybrid model alternating attention and Mamba-2 layers, as a must-read, and notes a 4B Nano variant and the 550B-A55B Nemotron 3 Ultra released two days earlier.
AICohere and Cohere Labs released North Mini Code 1.0, an open-weights 30B-A3B mixture-of-experts model for code generation and agentic terminal tasks, under Apache 2.0. The model has 256K context and 64K max output, and is trained for tool use. Its benchmark table lists Terminal-Bench v2 at 36.0, SWE-Bench Verified at 67.6, and LiveCodeBench v6 at 70.3, below Qwen3.6 on several tasks.
Why it matters: The card lists benchmark results against Qwen3.6, Gemma4, and other models, showing where North Mini Code trails on some coding and agentic tasks.
AIAntigravity has enabled the /teamwork-preview command for all paid plans, letting users run parallel implementation and verification agents on complex tasks. The post says the team built a working OS with it, but warns that it can consume a large number of tokens.
AIMaintainers and NVIDIA engineers improved multi-GPU performance in ggml, the low-level engine behind llama.cpp, yielding significant gains on RTX systems. The work also lays groundwork for hardware-agnostic tensor parallelism in ggml. Details are in a technical blog from NVIDIA RTX Spark.
AIBaidu CoBuddy is now available for free on Novita AI, a code-focused model aimed at developers and AI agents. It offers a 131K context window, up to 65K output tokens, and native tool calling. The model is served through Novita's serverless API for high-throughput, low-latency inference.
AICognition built an automated agent that classifies Devin sessions as productive and estimates the human engineering hours each one would have taken. On 233 held-out sessions the estimator reached an rlog of 0.74, with individual errors often 2 to 3 times in either direction but roughly unbiased in aggregate. The system is calibrated to underestimate and is currently running with Devin customers.
Why it matters: The post shows how the measurement design, from hours-based metrics to conservative calibration, determines whether agent productivity estimates can be trusted in aggregate.
AIVarun Mohan announced a new Antigravity IDE release that significantly reduces cases of memory usage blowing up. He said the team is pushing hard on performance and reliability and will address user feedback.
AIGoogle Labs has launched Dreambeans, an experimental mobile app that uses Personal Intelligence to connect to users' Google apps and deliver daily collections of personalized stories. The app surfaces relevant topics to help users explore what they care about most. It is available starting today to eligible US-based Google AI Ultra users aged 18 and older, with an open waitlist on the Google Labs website.
AIGeorgi Gerganov says Computex this year shows a strong signal for local AI, with major players like NVIDIA and Microsoft embracing and discussing local AI workloads. He adds that dedicated consumer hardware and models are on the way.
AIMiniMax published MiniMax-M3-MXFP8, an MXFP8 quantized variant of its native multimodal M3 model with 1M context, about 428B total parameters and about 23B activated parameters. M3 adds MiniMax Sparse Attention, which the source says yields 9× prefill and 15× decode speedups over M2 at 1M context. The model supports three thinking modes (enabled, adaptive, disabled) via the thinking parameter and can be served with SGLang, vLLM, or Transformers.
Why it matters: The release pairs sparse attention for 1M-token contexts with reported prefill and decode speedups over M2, useful for judging long-context serving costs.
AIXiaomi MiMo has published a blog detailing full-pipeline inference optimizations for the MiMo-V2.5 series. The post highlights how the team pushed hybrid sliding window attention (SWA) efficiency to its limit. The full write-up is available at
AIXiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention. The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations. It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.
Why it matters: The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.
AICognition describes autonomous testing in Devin, where the agent writes a source-grounded test plan, operates the app through computer use, and returns labeled screenshots and an annotated video. Login steps are handled by a deterministic testing skill, and the company says test runs approved per day more than doubled in recent months. Known limits include timing errors with transient UI elements and models sometimes triggering states through JavaScript instead of clicking the interface.
Why it matters: The post explains how computer use, test plans, deterministic login scripts, and annotated recordings let Devin verify its own code changes end to end.
AICognition raised over $1B at a $26B valuation, led by Lux Capital, General Catalyst, and 8VC. Enterprise Devin usage grew more than 10x since the start of this year, and run-rate revenue reached $492M. The company also said it launched SWE-1.6, which reaches up to 950 tok/s and has become the most used model in Devin Desktop.