Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Jul 29

Jul 29Wed
  1. Liquid AI NewsletterAI score46

    Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi

    AILiquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens. The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages. LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.

Jul 28

Jul 28Tue
  1. Fireworks AI BlogAI score46

    Fireworks AI Shows Low-Cost Fine-Tuning Lifts Domain Embedding Retrieval

    AIFireworks AI describes fine-tuning Qwen3-Embedding-8B on private (query, positive) pairs using bidirectional InfoNCE loss through its Training SDK, then serving the model via an OpenAI-compatible embeddings endpoint. The post reports that around 150 training steps was enough, that rank-32 LoRA landed within about one point of full-parameter fine-tuning, and that gains were largest where the base model struggled, while tasks like CoSQA and FiQA2018 showed flat results.

Jul 25

Jul 25Sat

Jul 24

Jul 24Fri

Jul 23

Jul 23Thu
  1. One Useful Thing (Ethan Mollick)AI score67

    Ethan Mollick's guide to choosing AI tools for agentic work

    AIEthan Mollick's guide says ChatGPT and Claude are the main choices for real work, since their agent modes can act on a computer. He separates agent modes that run on the company's computers from those that access the user's own computer. He recommends keeping approval settings on for sending, spending, or deleting, because of prompt injection risk. He also notes that Gemini currently lags for agentic work, though its Notebook and video tools are useful.

Jul 21

Jul 21Tue
  1. Andrej KarpathyAI score30

    Karpathy suggests long voice rambles help LLMs understand your intent

    AIAndrej Karpathy describes using /voice to ramble for about 10 minutes, sometimes as a short interview, to give an LLM context that would be tedious to type. He says LLMs reconstruct these messy streams of thought remarkably well, often returning a cleaner version than the speaker started with, which improves shared understanding and reduces later corrections.

Jul 19

Jul 19Sun

Jul 18

Jul 18Sat
  1. Ahead of AI (Sebastian Raschka)AI score52

    How Reasoning Effort Settings Are Built Into LLMs Through Training

    AIThe article explains how reasoning models can offer multiple effort modes, separating training-time methods from inference-time controls such as system prompts and chat templates. It compares six open-weight models, including DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5, GLM-5, Qwen3, and Inkling, noting that their reports disclose different levels of detail. It also shows how GPT-5.6's model selection and effort settings act as two separate scaling axes.

Jul 17

Jul 17Fri
  1. Andrew NgAI score28

    DeepLearning.AI launches course on fast LLM inference with Cerebras

    AIDeepLearning.AI has launched a short course, built with Cerebras, on building LLM applications that respond quickly using inference-optimized hardware. The course compares how GPUs, TPUs, and Cerebras' Wafer-Scale Engine handle the memory-to-compute bottleneck, which keeps model weights close to compute units to speed token generation. It covers real-time applications such as live translation and voice agents, plus habits for agentic coding.

    Video from @AndrewYNg's post

Jul 14

Jul 14Tue

Jul 12

Jul 12Sun
  1. OpenAI NewsroomAI score12

    James Costello uses ChatGPT to run his demolition business

    AIStructural engineer James Costello, who oversees complex New York City high-rise demolitions, uses ChatGPT to review lengthy contracts, organize compliance documents, and create construction plans for his family-rooted firm DEMTEC. The post says the tool helps him move through these workflows faster and with more confidence, freeing time to grow the business and support his team.

    Image from @OpenAINewsroom's post

Jul 11

Jul 11Sat
  1. OpenAI NewsroomAI score15

    Emma Dahl used ChatGPT to help design and build her custom wedding dress

    AIEmma Dahl wanted a historically inspired wedding dress that incorporated pearls from her grandmother's necklace, so she used ChatGPT over months to troubleshoot niche sewing and corset construction. The chatbot also helped her choose a sewing machine upgrade, the shape of her veil, and how to pack and travel with the orchids for her bouquet.

    Image from @OpenAINewsroom's post

Jul 6

Jul 6Mon

Jul 5

Jul 5Sun

Jun 30

Jun 30Tue
  1. Andrew NgAI score50

    Andrew Ng outlines three loops for building 0-to-1 AI products

    AIAndrew Ng describes three loops he uses to build 0-to-1 products with AI agents: an agentic coding loop, a developer feedback loop, and an external feedback loop. He says the agentic coding loop runs every few minutes, letting coding agents build, test, and iterate on software for around an hour without human intervention. The developer feedback loop operates over tens of minutes to hours, with humans steering product decisions because they hold a context advantage over AI systems.

    Image from @AndrewYNg's post

Jun 27

Jun 27Sat
  1. Ahead of AI (Sebastian Raschka)AI score37

    Local Coding Agents: Setting Up Qwen3.6 with Open-Source Harnesses

    AISebastian Raschka's tutorial shows how to build a fully local coding agent by pairing an open-weight LLM served through an inference runtime with an open-source harness that can read files, edit code, and run commands. He recommends Qwen-Code for Qwen3.6, citing Nvidia's Polar paper, which found Qwen models performed best in Qwen-Code. The Qwen3.6 35B-A3B model is about 22 GB to download and needs roughly 30–40 GB of RAM.

Jun 26

Jun 26Fri
  1. PaddlePaddleAI score32

    PP-OCRv6 Ep.4 benchmarks show 3.9x CPU speedup and 0.13s A100 OCR

    AIPaddlePaddle's PP-OCRv6 Tech Deep Dive Ep.4 benchmarks the OCR models across A100, V100, Intel Xeon CPU, and Apple M4 setups. PP-OCRv6_tiny processes an image in 0.13s on A100, while PP-OCRv6_tiny with OpenVINO runs 3.9x faster than PP-OCRv5_mobile on Intel CPU. The post recommends Medium for high-concurrency APIs, Small for CPU document systems, Tiny for mobile or embedded devices, and Medium or Small for multilingual business use.

    Image from @PaddlePaddle's post

Jun 25

Jun 25Thu
  1. Lilian WengAI score40

    Lilian Weng's Overview of Scaling Laws and Compute-Optimal Allocation

    AILilian Weng published a long blog post on scaling laws, which help estimate the best split of compute between data and model size before a large training run. The post covers what scaling laws predict, how compute-optimal allocation works, and why Kaplan et al. and Chinchilla reach different conclusions. It also addresses how data limits and fitting details make extrapolation difficult.

Jun 24

Jun 24Wed
  1. Eugene YanAI score33

    How benchmarks evaluate AI models' ability to find and exploit vulnerabilities

    AIThe post explains how cybersecurity benchmarks test whether models can find and exploit vulnerabilities. Common setups place a target in a sandboxed Docker container, provide either only code (0-day) or code plus a patch (1-day), allow tools like bash and static analyzers, and use a grader to score exploits or captured flags.

  2. PaddlePaddleAI score30

    PP-OCRv6 Detection Module Outperforms VLMs on Text Localization Benchmarks

    AIPaddlePaddle says its PP-OCRv6_medium text detector reached an 86.2% detection Hmean in benchmarks, versus 46.8% for Gemini-3.1-Pro and 38.3% for GPT-5.5. The detector's design uses RepLKFPN with 7×7 kernels to cut FPN neck parameters from 172K to 118K, auxiliary deep supervision heads on P2–P4, and Focal Loss paired with Dice Loss, which adds +1.15% Hmean in ablation.

    Image from @PaddlePaddle's post

Jun 22

Jun 22Mon
  1. Zed BlogAI score14

    Zed Blog's Hidden Gems Part 4 covers multi-project navigation and command aliases

    AIZed Blog's "Hidden Gems: Part 4" lists editor tips including a centered layout toggle with adjustable padding, keybindings for switching between projects, worktrees and branches, and setting EDITOR and VISUAL to zed --wait in the integrated terminal. It also explains command_aliases for mapping short mnemonics such as gd and gcp to commands like git::Diff and git::CreatePullRequest.

Jun 18

Jun 18Thu
  1. Andrew NgAI score15

    DeepLearning.AI launches course on adding voice to AI agents

    AIDeepLearning.AI has launched a course, taught by VocalBridge CEO Ashwyn, on adding voice to AI agents and applications. It covers building voice agents that are both reliable and fast, with three projects: a voice-interactive game, an agent that gains a voice in about 10 lines of code, and an agent that places outbound calls via a make_phone_call function.

    Video from @AndrewYNg's post

Jun 17

Jun 17Wed
  1. PromptArmor Threat IntelligenceAI score62

    PromptArmor shows Codex auto-review agent approved malware install via prompt injection

    AIPromptArmor demonstrated that OpenAI's Approve-for-me agent approved a malicious NPM install with elevated privileges after a hidden prompt injection in an external GitHub issue influenced the main Codex agent. The malicious package's post-install script then ran unsandboxed with the user's full privileges. The report also gives steps for organizations to disable agentic auto-review in Claude Code and Codex.

    Why it matters: The report shows a prompt-injected GitHub issue leading an approval agent to permit a malicious NPM install, a concrete test of agent-in-the-loop guardrails.

Jun 12

Jun 12Fri

Jun 9

Jun 9Tue

Jun 8

Jun 8Mon

May 30

May 30Sat
  1. Xiaomi MiMoAI score62

    Xiaomi details how it turned MiMo-V2.5 Hybrid SWA savings into production inference gains

    AIXiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention. The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations. It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.

    Why it matters: The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.

May 28

May 28Thu
  1. Cognition Blog (Devin, Windsurf)AI score62

    Devin Tests Its Own Code Changes in the Cloud and Returns Proof

    AICognition describes autonomous testing in Devin, where the agent writes a source-grounded test plan, operates the app through computer use, and returns labeled screenshots and an annotated video. Login steps are handled by a deterministic testing skill, and the company says test runs approved per day more than doubled in recent months. Known limits include timing errors with transient UI elements and models sometimes triggering states through JavaScript instead of clicking the interface.

    Why it matters: The post explains how computer use, test plans, deterministic login scripts, and annotated recordings let Devin verify its own code changes end to end.

May 25

May 25Mon
  1. MiniMax BlogAI score67

    MiniMax explains why its LLM failed to generate the name Ma Jiaqi

    AIMiniMax says its M2 series could not output the name Ma Jiaqi, a failure it traced to post-training data that rarely included the token. Its tests found the input embedding stayed stable while the lm_head weights for low-frequency tokens drifted during SFT. A synthetic full-vocabulary repetition dataset restored generation for affected tokens and reduced Japanese-to-Russian confusion from 47% to 1%.

    Why it matters: The post traces a community-noticed token failure through tokenizer, embedding, and lm_head tests, showing how post-training data coverage can cause low-frequency token drift.

May 24

May 24Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score45

    Fun-ASR-Nano-2512-hf: Alibaba's Speech Recognition Model Gets Transformers Version

    AIFunAudioLLM has released Fun-ASR-Nano-2512-hf, a Hugging Face Transformers-compatible version of its end-to-end speech recognition model, which supports Chinese, English, and Japanese. The Chinese coverage includes 7 dialect groups and 26 regional accents, and a separate Fun-ASR-MLT-Nano-2512 checkpoint handles 31-language recognition. Developers can run the model natively in Transformers 5.17.0 without custom model code or trust_remote_code=True.

May 21

May 21Thu
  1. Tri DaoAI score44

    Transformers reduce to GEMM-plus-epilogue, enabling LLM-written fast kernels

    AITri Dao says that after a mathematical rewrite, all transformer operations can be expressed as a series of GEMMs with epilogues. Given a few optimized primitives, LLMs and novice humans can write near speed-of-light kernels for transformer ops. The related CODA work fuses memory-bound surrounding ops into the matmul epilogue, and LLMs can also write CODA kernels approaching speed-of-light.

May 18

May 18Mon
  1. Eugene YanAI score62

    Cloudflare outlines an eight-stage agent harness for vulnerability discovery

    AIEugene Yan shares Cloudflare's description of a vulnerability discovery harness that runs eight stages, from reconnaissance to report writing. The pipeline uses about 50 concurrent agents to hunt for bugs, independent agents to try to disprove findings, and a trace step to confirm whether attacker input reaches each bug. Reachable findings feed back into new hunt tasks before a report is written against a predefined schema.

    Image from @eugeneyan's post

May 16

May 16Sat
  1. Ahead of AI (Sebastian Raschka)AI score62

    Recent LLM architecture changes that cut long-context KV cache and attention cost

    AISebastian Raschka reviews recent open-weight LLM architecture changes aimed at reducing long-context memory and compute costs. He covers KV sharing and per-layer embeddings in Gemma 4, per-layer query-head budgeting in Laguna XS.2, Compressed Convolutional Attention in ZAYA1-8B, and mHC with CSA/HCA compressed attention in DeepSeek V4. The article reports that DeepSeek V4-Pro uses 27% of single-token inference FLOPs and 10% of the KV cache size of DeepSeek V3.2 at a 1M-token context.