Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Jul 29

Jul 29Wed
  1. Liquid AI NewsletterAI score46

    Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi

    AILiquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens. The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages. LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.

Jul 28

Jul 28Tue
  1. Fireworks AI BlogAI score46

    Fireworks AI Shows Low-Cost Fine-Tuning Lifts Domain Embedding Retrieval

    AIFireworks AI describes fine-tuning Qwen3-Embedding-8B on private (query, positive) pairs using bidirectional InfoNCE loss through its Training SDK, then serving the model via an OpenAI-compatible embeddings endpoint. The post reports that around 150 training steps was enough, that rank-32 LoRA landed within about one point of full-parameter fine-tuning, and that gains were largest where the base model struggled, while tasks like CoSQA and FiQA2018 showed flat results.

Jul 25

Jul 25Sat

Jul 24

Jul 24Fri

Jul 23

Jul 23Thu
  1. One Useful Thing (Ethan Mollick)AI score67

    Ethan Mollick's guide to choosing AI tools for agentic work

    AIEthan Mollick's guide says ChatGPT and Claude are the main choices for real work, since their agent modes can act on a computer. He separates agent modes that run on the company's computers from those that access the user's own computer. He recommends keeping approval settings on for sending, spending, or deleting, because of prompt injection risk. He also notes that Gemini currently lags for agentic work, though its Notebook and video tools are useful.

Jul 21

Jul 21Tue
  1. Andrej KarpathyAI score30

    Karpathy suggests long voice rambles help LLMs understand your intent

    AIAndrej Karpathy describes using /voice to ramble for about 10 minutes, sometimes as a short interview, to give an LLM context that would be tedious to type. He says LLMs reconstruct these messy streams of thought remarkably well, often returning a cleaner version than the speaker started with, which improves shared understanding and reduces later corrections.

Jul 18

Jul 18Sat
  1. Ahead of AI (Sebastian Raschka)AI score52

    How Reasoning Effort Settings Are Built Into LLMs Through Training

    AIThe article explains how reasoning models can offer multiple effort modes, separating training-time methods from inference-time controls such as system prompts and chat templates. It compares six open-weight models, including DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5, GLM-5, Qwen3, and Inkling, noting that their reports disclose different levels of detail. It also shows how GPT-5.6's model selection and effort settings act as two separate scaling axes.

Jul 17

Jul 17Fri
  1. Andrew NgAI score28

    DeepLearning.AI launches course on fast LLM inference with Cerebras

    AIDeepLearning.AI has launched a short course, built with Cerebras, on building LLM applications that respond quickly using inference-optimized hardware. The course compares how GPUs, TPUs, and Cerebras' Wafer-Scale Engine handle the memory-to-compute bottleneck, which keeps model weights close to compute units to speed token generation. It covers real-time applications such as live translation and voice agents, plus habits for agentic coding.

Jul 6

Jul 6Mon

Jun 30

Jun 30Tue
  1. Andrew NgAI score50

    Andrew Ng outlines three loops for building 0-to-1 AI products

    AIAndrew Ng describes three loops he uses to build 0-to-1 products with AI agents: an agentic coding loop, a developer feedback loop, and an external feedback loop. He says the agentic coding loop runs every few minutes, letting coding agents build, test, and iterate on software for around an hour without human intervention. The developer feedback loop operates over tens of minutes to hours, with humans steering product decisions because they hold a context advantage over AI systems.

Jun 27

Jun 27Sat
  1. Ahead of AI (Sebastian Raschka)AI score37

    Local Coding Agents: Setting Up Qwen3.6 with Open-Source Harnesses

    AISebastian Raschka's tutorial shows how to build a fully local coding agent by pairing an open-weight LLM served through an inference runtime with an open-source harness that can read files, edit code, and run commands. He recommends Qwen-Code for Qwen3.6, citing Nvidia's Polar paper, which found Qwen models performed best in Qwen-Code. The Qwen3.6 35B-A3B model is about 22 GB to download and needs roughly 30–40 GB of RAM.

Jun 26

Jun 26Fri
  1. PaddlePaddleAI score32

    PP-OCRv6 Ep.4 benchmarks show 3.9x CPU speedup and 0.13s A100 OCR

    AIPaddlePaddle's PP-OCRv6 Tech Deep Dive Ep.4 benchmarks the OCR models across A100, V100, Intel Xeon CPU, and Apple M4 setups. PP-OCRv6_tiny processes an image in 0.13s on A100, while PP-OCRv6_tiny with OpenVINO runs 3.9x faster than PP-OCRv5_mobile on Intel CPU. The post recommends Medium for high-concurrency APIs, Small for CPU document systems, Tiny for mobile or embedded devices, and Medium or Small for multilingual business use.

Jun 25

Jun 25Thu
  1. Lilian WengAI score40

    Lilian Weng's Overview of Scaling Laws and Compute-Optimal Allocation

    AILilian Weng published a long blog post on scaling laws, which help estimate the best split of compute between data and model size before a large training run. The post covers what scaling laws predict, how compute-optimal allocation works, and why Kaplan et al. and Chinchilla reach different conclusions. It also addresses how data limits and fitting details make extrapolation difficult.

Jun 24

Jun 24Wed
  1. PaddlePaddleAI score30

    PP-OCRv6 Detection Module Outperforms VLMs on Text Localization Benchmarks

    AIPaddlePaddle says its PP-OCRv6_medium text detector reached an 86.2% detection Hmean in benchmarks, versus 46.8% for Gemini-3.1-Pro and 38.3% for GPT-5.5. The detector's design uses RepLKFPN with 7×7 kernels to cut FPN neck parameters from 172K to 118K, auxiliary deep supervision heads on P2–P4, and Focal Loss paired with Dice Loss, which adds +1.15% Hmean in ablation.

Jun 17

Jun 17Wed
  1. PromptArmor Threat IntelligenceAI score62

    PromptArmor shows Codex auto-review agent approved malware install via prompt injection

    AIPromptArmor demonstrated that OpenAI's Approve-for-me agent approved a malicious NPM install with elevated privileges after a hidden prompt injection in an external GitHub issue influenced the main Codex agent. The malicious package's post-install script then ran unsandboxed with the user's full privileges. The report also gives steps for organizations to disable agentic auto-review in Claude Code and Codex.

    Why it matters: The report shows a prompt-injected GitHub issue leading an approval agent to permit a malicious NPM install, a concrete test of agent-in-the-loop guardrails.

Jun 12

Jun 12Fri

Jun 9

Jun 9Tue

Jun 8

Jun 8Mon

May 30

May 30Sat
  1. Xiaomi MiMoAI score62

    Xiaomi details how it turned MiMo-V2.5 Hybrid SWA savings into production inference gains

    AIXiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention. The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations. It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.

    Why it matters: The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.

May 28

May 28Thu
  1. Cognition Blog (Devin, Windsurf)AI score62

    Devin Tests Its Own Code Changes in the Cloud and Returns Proof

    AICognition describes autonomous testing in Devin, where the agent writes a source-grounded test plan, operates the app through computer use, and returns labeled screenshots and an annotated video. Login steps are handled by a deterministic testing skill, and the company says test runs approved per day more than doubled in recent months. Known limits include timing errors with transient UI elements and models sometimes triggering states through JavaScript instead of clicking the interface.

    Why it matters: The post explains how computer use, test plans, deterministic login scripts, and annotated recordings let Devin verify its own code changes end to end.

May 25

May 25Mon
  1. MiniMax BlogAI score67

    MiniMax explains why its LLM failed to generate the name Ma Jiaqi

    AIMiniMax says its M2 series could not output the name Ma Jiaqi, a failure it traced to post-training data that rarely included the token. Its tests found the input embedding stayed stable while the lm_head weights for low-frequency tokens drifted during SFT. A synthetic full-vocabulary repetition dataset restored generation for affected tokens and reduced Japanese-to-Russian confusion from 47% to 1%.

    Why it matters: The post traces a community-noticed token failure through tokenizer, embedding, and lm_head tests, showing how post-training data coverage can cause low-frequency token drift.

May 24

May 24Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score45

    Fun-ASR-Nano-2512-hf: Alibaba's Speech Recognition Model Gets Transformers Version

    AIFunAudioLLM has released Fun-ASR-Nano-2512-hf, a Hugging Face Transformers-compatible version of its end-to-end speech recognition model, which supports Chinese, English, and Japanese. The Chinese coverage includes 7 dialect groups and 26 regional accents, and a separate Fun-ASR-MLT-Nano-2512 checkpoint handles 31-language recognition. Developers can run the model natively in Transformers 5.17.0 without custom model code or trust_remote_code=True.

May 21

May 21Thu
  1. Tri DaoAI score44

    Transformers reduce to GEMM-plus-epilogue, enabling LLM-written fast kernels

    AITri Dao says that after a mathematical rewrite, all transformer operations can be expressed as a series of GEMMs with epilogues. Given a few optimized primitives, LLMs and novice humans can write near speed-of-light kernels for transformer ops. The related CODA work fuses memory-bound surrounding ops into the matmul epilogue, and LLMs can also write CODA kernels approaching speed-of-light.

May 18

May 18Mon
  1. Eugene YanAI score62

    Cloudflare outlines an eight-stage agent harness for vulnerability discovery

    AIEugene Yan shares Cloudflare's description of a vulnerability discovery harness that runs eight stages, from reconnaissance to report writing. The pipeline uses about 50 concurrent agents to hunt for bugs, independent agents to try to disprove findings, and a trace step to confirm whether attacker input reaches each bug. Reachable findings feed back into new hunt tasks before a report is written against a predefined schema.

May 16

May 16Sat
  1. Ahead of AI (Sebastian Raschka)AI score62

    Recent LLM architecture changes that cut long-context KV cache and attention cost

    AISebastian Raschka reviews recent open-weight LLM architecture changes aimed at reducing long-context memory and compute costs. He covers KV sharing and per-layer embeddings in Gemma 4, per-layer query-head budgeting in Laguna XS.2, Compressed Convolutional Attention in ZAYA1-8B, and mHC with CSA/HCA compressed attention in DeepSeek V4. The article reports that DeepSeek V4-Pro uses 27% of single-token inference FLOPs and 10% of the KV cache size of DeepSeek V3.2 at a 1M-token context.

May 10

May 10Sun
  1. Cognition Blog (Devin, Windsurf)AI score39

    Devin Automates HIL/SIL Failure Triage and Scales Test Generation at Automotive Firms

    AICognition reports that deploying its Devin agent on hardware-in-the-loop and software-in-the-loop workflows cut failure triage time and multiplied test generation at automotive customers. One team reclaimed 2K–4K engineering hours monthly across about 4,000 tickets, while RV Tech rose from 1–2 to 10–15 generated tests per day. Devin also helps convert bottlenecked HIL tests into SIL equivalents to catch failures earlier.

Apr 23

Apr 23Thu

Apr 22

Apr 22Wed

Apr 21

Apr 21Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition says multi-agent systems work when only one agent writes

    AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.

    Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.

Apr 17

Apr 17Fri

Apr 16

Apr 16Thu

Apr 4

Apr 4Sat
  1. Andrej KarpathyAI score62

    Andrej Karpathy outlines an LLM-maintained markdown wiki workflow for personal research

    AIKarpathy describes using LLMs to compile raw source documents into a markdown wiki that he views in Obsidian, with the LLM writing and maintaining most of the wiki. He reports that at about 100 articles and 400K words, the LLM agent can answer complex questions directly from the wiki, and he also runs LLM health checks to find inconsistencies and gaps. He shares the underlying idea as an "idea file" that users can give to their own agents to build a customized version.

Apr 2

Apr 2Thu
  1. Andrej KarpathyAI score49

    Karpathy shares an LLM-maintained personal knowledge base workflow

    AIAndrej Karpathy describes using LLMs to compile raw research sources into a markdown wiki of about 100 articles and 400K words, viewed in Obsidian. He says an LLM agent answers complex questions against the wiki without RAG, with outputs filed back to enhance it. He also suggests the workflow could become a product rather than a collection of scripts.

Mar 24

Mar 24Tue
  1. Anthropic EngineeringAI score78

    How Anthropic built Claude Code auto mode to replace skipped permissions

    AIAnthropic describes Claude Code auto mode, which delegates approval of agent actions to model-based classifiers instead of manual prompts or skipped permissions. The classifier reviews tool calls before execution and a separate probe screens tool outputs for prompt injection. Anthropic reports a 0.4% false positive rate on real internal traffic and a 17% false negative rate on real overeager actions.

    Why it matters: The post explains the layered classifier design and its measured tradeoffs, showing how autonomous coding agents can cut approval fatigue without fully removing risk.