Midjourney unveils technical dive into its new Scanner
AIA technical dive inside our new "Midjourney Scanner"
Updated
Updated
AIA technical dive inside our new "Midjourney Scanner"
AIPromptArmor demonstrated that OpenAI's Approve-for-me agent approved a malicious NPM install with elevated privileges after a hidden prompt injection in an external GitHub issue influenced the main Codex agent. The malicious package's post-install script then ran unsandboxed with the user's full privileges. The report also gives steps for organizations to disable agentic auto-review in Claude Code and Codex.
Why it matters: The report shows a prompt-injected GitHub issue leading an approval agent to permit a malicious NPM install, a concrete test of agent-in-the-loop guardrails.
AIAwni Hannun praised a WWDC video by @angeloskath explaining how to set up agentic AI to run locally with MLX in an approachable way. He said the demos work well now, which was impractical less than a year ago before M5 and recent gains in open-weights models.
AIStability AI says Stable Audio 3.0 was built to support exploratory audio work such as stem remixing. A quoted post from @teropa reports using the Medium model with an init_audio input and an init_noise_level of 0.4–0.5, with empty prompts.
AIApple's MLX team released three videos at WWDC covering running agents locally, distributed inference and training, and MLX Swift. The post shares YouTube links for each talk, presented by Angelos Katharopoulos, Tatiana Likhomanenko, and David Koski.
AIMiniMax describes MaxProof, an evolutionary-search framework that lets its M3 model refine candidate proofs over multiple rounds. The post says the M3 model exceeded the human gold-medal threshold on the IMO 2025 and USAMO 2026 benchmarks with MaxProof, and explains the Proof RL, verifier alignment, and refinement training behind it.
AIXiaomi MiMo has published a blog detailing full-pipeline inference optimizations for the MiMo-V2.5 series. The post highlights how the team pushed hybrid sliding window attention (SWA) efficiency to its limit. The full write-up is available at
AIXiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention. The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations. It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.
Why it matters: The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.
AICognition describes autonomous testing in Devin, where the agent writes a source-grounded test plan, operates the app through computer use, and returns labeled screenshots and an annotated video. Login steps are handled by a deterministic testing skill, and the company says test runs approved per day more than doubled in recent months. Known limits include timing errors with transient UI elements and models sometimes triggering states through JavaScript instead of clicking the interface.
Why it matters: The post explains how computer use, test plans, deterministic login scripts, and annotated recordings let Devin verify its own code changes end to end.
AIMiniMax says the M2 series could not generate the rare token "嘉祺" in names like Ma Jiaqi, and its investigation traced the cause to post-training. The company found the token was learned in pretraining, but low coverage of rare tokens in post-training data caused lm_head vectors to drift. Adding synthetic full-vocabulary repetition data restored generation for these tokens and reduced Japanese-to-Russian mixing from 47% to 1%.
Why it matters: The post traces a specific token failure through tokenizer, embedding, and lm_head checks, showing a reusable way to diagnose post-training generation problems.
AIFunAudioLLM has released Fun-ASR-Nano-2512-hf, a Hugging Face Transformers-compatible version of its end-to-end speech recognition model, which supports Chinese, English, and Japanese. The Chinese coverage includes 7 dialect groups and 26 regional accents, and a separate Fun-ASR-MLT-Nano-2512 checkpoint handles 31-language recognition. Developers can run the model natively in Transformers 5.17.0 without custom model code or trust_remote_code=True.
AITri Dao says that after a mathematical rewrite, all transformer operations can be expressed as a series of GEMMs with epilogues. Given a few optimized primitives, LLMs and novice humans can write near speed-of-light kernels for transformer ops. The related CODA work fuses memory-bound surrounding ops into the matmul epilogue, and LLMs can also write CODA kernels approaching speed-of-light.
AIEugene Yan shares Cloudflare's description of a vulnerability discovery harness that runs eight stages, from reconnaissance to report writing. The pipeline uses about 50 concurrent agents to hunt for bugs, independent agents to try to disprove findings, and a trace step to confirm whether attacker input reaches each bug. Reachable findings feed back into new hunt tasks before a report is written against a predefined schema.
AISebastian Raschka reviews recent open-weight LLM architecture changes aimed at reducing long-context memory and compute costs. He covers KV sharing and per-layer embeddings in Gemma 4, per-layer query-head budgeting in Laguna XS.2, Compressed Convolutional Attention in ZAYA1-8B, and mHC with CSA/HCA compressed attention in DeepSeek V4. The article reports that DeepSeek V4-Pro uses 27% of single-token inference FLOPs and 10% of the KV cache size of DeepSeek V3.2 at a 1M-token context.
AICognition reports that deploying its Devin agent on hardware-in-the-loop and software-in-the-loop workflows cut failure triage time and multiplied test generation at automotive customers. One team reclaimed 2K–4K engineering hours monthly across about 4,000 tickets, while RV Tech rose from 1–2 to 10–15 generated tests per day. Devin also helps convert bottlenecked HIL tests into SIL equivalents to catch failures earlier.
AIAwni Hannun suggests stating your own understanding of a topic before asking an AI, then having the model correct your mistakes. He argues this approach is a much better way to learn than asking the question directly.
AINVIDIA and Google DeepMind experts will host a live session on Friday, April 24 at 11:00 AM PDT showcasing Gemma 4 running on DGX Spark. The demos cover vision translation, long-context document Q&A, and real-time code generation, with audience questions welcome.
AITri Dao points to Wentao Guo's explanation of a mathematical rewrite of the MoE backward pass, which reduces activation memory and speeds up training, especially for fine-grained MoE. He also notes the work leverages new Blackwell features, including 2CTA MMA and CLC, to build fast MoE kernels.
AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.
Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.
AINVIDIA AI Developer announced a live broadcast on contributing to open source, featuring NVIDIA NemoClaw and Nemotron Labs. The post links only to the broadcast and gives no further details on the content or speakers.
AIA developer trained a 12M-parameter LLM on a custom ML framework built with a Rust backend and CUDA kernels, including Flash Attention, fused LayerNorm, and fused GELU. The framework claims 3x throughput gains, WebGPU fallback for non-NVIDIA devices, and a TypeScript API installable via npm.
AINVIDIA's developer account published a step-by-step tutorial for building a fully local, sandboxed, always-on AI agent. The guide uses OpenClaw with NVIDIA NemoClaw and runs on NVIDIA DGX Spark.
AIA blog post by @pcuenq describes a new skill and test harness for automating the porting of new models from Transformers to mlx-lm. The post is shared by Awni Hannun, who highlights it as a useful guide for this workflow.
AIKarpathy describes using LLMs to compile raw source documents into a markdown wiki that he views in Obsidian, with the LLM writing and maintaining most of the wiki. He reports that at about 100 articles and 400K words, the LLM agent can answer complex questions directly from the wiki, and he also runs LLM health checks to find inconsistencies and gaps. He shares the underlying idea as an "idea file" that users can give to their own agents to build a customized version.
AIAndrej Karpathy describes using LLMs to compile raw research sources into a markdown wiki of about 100 articles and 400K words, viewed in Obsidian. He says an LLM agent answers complex questions against the wiki without RAG, with outputs filed back to enhance it. He also suggests the workflow could become a product rather than a collection of scripts.
AIAnthropic describes Claude Code auto mode, which delegates approval of agent actions to model-based classifiers instead of manual prompts or skipped permissions. The classifier reviews tool calls before execution and a separate probe screens tool outputs for prompt injection. Anthropic reports a 0.4% false positive rate on real internal traffic and a 17% false negative rate on real overeager actions.
Why it matters: The post explains the layered classifier design and its measured tradeoffs, showing how autonomous coding agents can cut approval fatigue without fully removing risk.
AIAnthropic's Labs team describes a three-agent harness with planner, generator, and evaluator agents for building full-stack applications over multi-hour autonomous coding sessions. The evaluator uses Playwright to test the running app against sprint contracts, and a retro game maker built with the harness worked end to end where a single-agent run's core feature did not. The author later removed the sprint construct and kept only the components still needed on Opus 4.6.
Why it matters: The post shows how a generator-evaluator loop, with explicit grading criteria and a tuned QA agent, turned a solo run's broken output into a working app, and how the harness was pruned as models improved.
AIPublic AI benchmarks such as MMLU, SWE-bench Verified, and GPQA Diamond are saturating or showing contamination, prompting OpenAI to call SWE-bench Verified "no longer suitable" in late February and recommend SWE-bench Pro. OpenAI's audit found 59.4% of the problems its best model failed had flawed test cases, and GPT-5.2, Claude Opus 4.5, and Gemini 3 Flash could reproduce original fixes from memory. The article argues that behavioral tests, such as Vending-Bench's simulated vending machine business, may be more useful for everyday model choice.
AIMax Woolf published a blog post summarizing what he learned over several months working with agents including Opus 4.5 and beyond. The post also describes an optimization trick he considers promising.
AICognition reports merging 659 Devin PRs into its own codebase last week, up from 154 in its best week in 2025. The post describes internal workflows across web, Slack, Linear, CLI, and API, including Devin Review for PR diffs and bug catching, a daily design system audit, automated bug triage on Linear, and DANA for data analysis.
Why it matters: The post shows concrete workflows for using Devin across Slack, Linear, and code review, with specific usage figures that help teams judge fit for their own engineering processes.
AIThe article argues that engineers are splitting their work between building a harness of constraints, tools, and documentation for agents and directing the agents' work. It cites OpenAI, Stripe, and Anthropic examples, including architecture guardrails, custom linter messages, AGENTS.md updates, and plan-first execution. The author notes that open problems remain around code maintainability, verification at scale, and adopting these practices in older codebases.
AIAnthropic's Claude API now supports automatic prefix caching, enabled by setting one cache control field at the top level of the request body. Developers still need to structure prompt templates so the reusable prefix stays consistent to benefit from caching.
AIOriol Vinyals shared a prompt asking an AI model to generate an SVG of a pelican riding a car in France, with a cat beside it and the Eiffel Tower in the background. The post offers no model name, results, or benchmark figures.
AIGenerate an SVG of a rollercoaster
AIMax Woolf reports that experiments using GPT-5.3-Codex to hyperoptimize batch prediction are going well. The post gives no benchmark scores, speedups, or further technical details.
AINicholas Carlini of Anthropic's Safeguards team describes an agent-team setup where 16 Claude instances worked in parallel on a shared codebase without human intervention to write a Rust-based C compiler. Over nearly 2,000 Claude Code sessions costing about $20,000 in API fees, the team produced a 100,000-line compiler that can build Linux 6.9 on x86, ARM, and RISC-V. The post focuses on harness design, including high-quality tests, lock files for task claiming, GCC as a reference oracle for the kernel, and the limits the project reached.
Why it matters: The post shows concrete harness design choices for long-running agent teams, including test design, locking, and parallel work division, that readers can adapt to their own autonomous projects.
AIAnthropic's performance engineering lead Tristan Hume describes how a take-home test for hiring performance engineers was repeatedly defeated by successive Claude models. Claude Opus 4 outperformed most human applicants within the 4-hour limit, and Claude Opus 4.5 matched the best candidates in 2 hours. Anthropic is releasing the original take-home as an open challenge, with the best known Claude result at 1487 cycles.
Why it matters: The post traces how each Claude model defeated the take-home test, showing concrete redesign tradeoffs for evaluating engineers when AI assistance is available.
AITim Dettmers, a professor who has used Claude Code for eight months to automate his own work, argues that more than 90% of code and text should be written by agents. He says the coding-focused hype on Twitter overstates parallel sessions and autonomy, which translate poorly to most non-software tasks. The post offers a balanced guide to what actually works in agent-based automation.
AICognition says its autonomous coding agent Devin can complete a .NET Framework to .NET Core migration in as little as two weeks, using a Strangler Fig approach adapted from Jimmy Bogard's guide. The post says Devin handles planning via Ask Devin and DeepWiki, dependency sharing, controller and view conversion, and session state adaptation through a remote app.
AICognition rebuilt its Devin coding agent for Claude Sonnet 4.5, reporting 2x faster performance and 12% better results on its Junior Developer Evals, now available in Agent Preview. The team found the model is aware of its context window, which led to premature wrap-up behavior that they countered with repeated prompts and a 200k usage cap within a 1M token beta.
Why it matters: The post explains which agent behaviors changed under Sonnet 4.5, such as context-window awareness and note-taking, that forced a rebuild rather than a simple model swap.