Skip to contentSkip to stories

Updated

#Tutorial/How-to

Showing low-relevance items too. Hide low-relevance items

Sep 23

Sep 23Wed
  1. Simon WillisonXAI score37

    Gemini 3.8 TTS models add cheap multi-voice conversation generation

    AIGoogle's new Gemini 3.8 TTS models generate conversations between multiple voices, with over 2,000 preset voices available and voice cloning supported. Simon Willison says the models are super-cheap and built a small UI to try them, with Claude writing a script in which two pelicans debate moving to Pacifica Pier.

  2. Boris ChernyXAI score51

    Anthropic details how it made claude.ai 3x faster in two weeks

    AIAnthropic says it made claude.ai 3x faster in two weeks, and the same speedup applies to the Desktop app. A linked blog post explains how the team used Claude to measure, debug, and improve performance, and includes prompts and methods.

  3. eric zakariassonXAI score67

    Cursor shares a prompt for reducing token cost in agent harnesses

    AICursor's Eric Zakariasson shared a prompt for improving an LLM agent harness to lower token cost per completed task without losing quality. The prompt covers the system prompt, tool definitions, cache layout, tool results, compaction, and subagents, and reports that one team's round of these changes cut overall token cost about 7%.

    Why it matters: The prompt gives a concrete checklist for cutting agent token cost per completed task, with tested figures on cache layout, tool offloading, and compaction.

  4. InferactOfficialAI score44

    vLLM maintainers show TPUv7 megakernels beat GB200 NVL72 on Kimi K3

    AIInferact says vLLM maintainers used megakernel optimization to reach 700 tokens per second per user on TPUv7 running Kimi K3. SemiAnalysis, which shared the work, reports this is 56% better performance than Nvidia's GB200 NVL72. Inferact links a full technical breakdown of the TPU megakernel work on its blog.

  5. GitHub Blog · AI & MLOfficialAI score46

    Copilot app rebuilds pull request view to render a 2,200-file diff smoothly

    AIGitHub rebuilt the pull request view in the GitHub Copilot app to keep review fast on very large diffs, testing it on an open source pull request with 2,200 files, over a million changed lines, and more than 400 inline review comments. The core difficulty is that review comment heights can only be measured at render time, which breaks the fixed-geometry virtualization used for code-only diffs. GitHub split the document height into a deterministic code domain and a separately measured domain for comment blocks.

  6. eric zakariassonXAI score36

    Optimizing reading for AI agents cuts context-gathering costs

    AIEric Zakariasson argues that agents spend heavily on reading context before and after work, so optimizing that reading makes a major difference. He recommends the linked guide to builders, or handing it to an agent to implement its findings. Cursor's related post reports 7% lower token costs with no drop in agent quality, achieved through tighter prompts, selective tool loading, better caching, and compressed file reads.

    Image from @ericzakariasson's post
  7. Microsoft ResearchOfficialAI score60

    Microsoft Research shows offloading robot AI inference improves performance and battery life

    AIMicrosoft Research reports that running physical AI inference on onboard GPUs can limit robot performance and battery life, while offloading inference to edge or cloud GPUs improved results in mobile manipulation tests. In its evaluation, smaller onboard GPUs slowed mapping and planning by up to 383% compared with an A100, and large onboard GPUs such as Jetson Thor drained robot batteries by up to 160%.

    Why it matters: The study measures how offloading robot inference to edge or cloud GPUs changes task success, battery life, and model size, offering evidence for infrastructure design.

Sep 22

Sep 22Tue
  1. TinkerOfficialAI score25

    Tinker fine-tunes Qwen3.6 for Jev-style probability prompts in 10 minutes

    AITinker says an open LLM can serve a Jev-like interface that takes discrete options and returns fast probabilities, since next-token prediction is already a probabilistic classifier. A post by @ekzhang1 reports that a $5, 10-minute supervised fine-tuning run on Tinker improved Qwen3.6-35B-A3B's handling of Jev-style prompts, with +8% on GPQA Diamond and +12% on MMLU-Pro.

  2. Google Developers BlogOfficialAI score62

    Antigravity SDK adds local Gemma 4 26B agent support via LiteRT

    AIGoogle announced that the Antigravity SDK supports local agent workflows, with initial support for Gemma 4 26B A4B through Google AI Edge's LiteRT. The post includes Python setup steps and says a recommended machine has more than 24GB VRAM or unified memory. It also describes a hybrid pattern in which a cloud Gemini 3.8 Flash planner hands work to local Gemma 4 26B models, with 97.2% of tokens in one recorded run staying local.

    Why it matters: The source shows how to run an agent with a local Gemma 4 26B model using LiteRT, plus a hybrid cloud-planner pattern that keeps most tokens on-device.

  3. Together AI BlogOfficialAI score38

    How to train your own Jev classifier for $17 with Together AI

    AIThe Together AI blog shows how to fine-tune a Qwen3.5 4B base model into a classification model using about 38,000 examples sampled from six Hugging Face datasets, at a training cost of roughly $17.0. The tutorial covers cloning the tev1 repository, normalizing data with provided scripts, launching a Together AI fine-tuning job that takes about 25 minutes, and deploying the result to a dedicated H100 endpoint.

  4. Boris ChernyXAI score42

    Boris Cherny Uses Opus 5.5 to Formally Verify Claude Agent SDK

    AIBoris Cherny used Opus 5.5 to formally verify the Claude Agent SDK with Lean, and short prompts produced 16 PRs fixing bugs and race conditions. He also combines Lean and TLA+ to find issues in data flow, concurrency, and state management, and says Claude is strong in both languages even though he does not know them well.

    Video from @bcherny's post
  5. Alex AlbertXAI score37

    Claude prompt recreates 1906 Market Street in Blender for video

    AIA prompt shared by Alex Albert asks Claude to recreate San Francisco's Market Street as it stood on April 17, 1906, before the earthquake, using Blender. It requires building a source file from Sanborn fire insurance maps, the Miles Brothers film, period photos, and USGS topography, with reusable Blender Python generators for facades, street lamps, and vehicles, ending in a 10-second video up the street.

  6. Alex AlbertXAI score18

    Opus 5.5 builds a 1906 San Francisco street scene in Blender

    AIAlex Albert says Opus 5.5 has improved 3D modeling and vision for Blender work, letting users build an entire world from a single prompt. He shares a historically accurate render of San Francisco's Market Street in 1906, before the earthquake.

    Video from @alexalbert__'s post
  7. Unsloth AIOfficialAI score70

    Qwen-Image-2.1 runs locally on 12GB VRAM using Unsloth GGUFs

    AIUnsloth says the 7B Qwen-Image-2.1 text-to-image and editing model can run locally on 12GB VRAM using its GGUF builds. It also states that the model performs on par with Nano Banana 2.0, and that Dynamic FP8 can run on 6GB of VRAM via offloading for higher quality. The image lists int8 at 7.26 GB with mean LPIPS 0.064 and fp8 at 7.12 GB with mean LPIPS 0.112, and says int8 is the default.

    Why it matters: The post gives concrete local-run settings, VRAM figures, and GGUF and FP8 options, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  8. Kimi.aiOfficialAI score46

    Kimi launches browser extension for chatting, automating web tasks

    AIKimi has released its Kimi Browser Extension, formerly Kimi WebBridge, which runs in the browser sidebar to navigate websites and fill out forms. Users can record repetitive steps once and save them as a skill for Kimi to reuse later. The extension is available now on the Chrome Web Store.

    Video from @Kimi_Moonshot's post
  9. OpenBMBOfficialAI score20

    OpenBMB praises MiniCPM5-2B workers in multi-agent invoice reconciliation

    AIOpenBMB thanked a developer for testing MiniCPM5-2B as a worker in a multi-agent workflow handling invoice matching, short payments, duplicate references, and disputes through tool calls. The background post says GPT-6 Astra coordinated the MiniCPM5-2B workers, verifying 32 synthetic invoices in 67.8 seconds with 232 executed tool calls. The demo does not move money.

  10. Suno BlogOfficialAI score16

    Suno Studio's Compressor Reduces Dynamic Range to Balance Mixes

    AISuno Studio includes a compressor plugin that reduces dynamic range by turning down peaks and applying makeup gain to raise quieter parts. Its controls include threshold, ratio, attack, and release, and it is especially useful for shaping uploaded or recorded audio. Studio is browser-based and comes free with every Suno Premier subscription.

Sep 21

Sep 21Mon
  1. Tencent HyOfficialAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

  2. xAI News (Grok)OfficialAI score46

    How SpaceXAI uses Grok Bot to scale customer support without new hires

    AISpaceXAI says its combined support team handled a 175% rise in tickets without hiring, crediting Grok Bot, which it says would otherwise have required about 200 additional staff. The company reports resolving tickets for $0.20 to $0.30 each, versus the $1 to $4 per resolution it attributes to traditional AI support tools. Grok Bot is also reported to resolve 99% of refund requests without human intervention.

  3. Together AI BlogOfficialAI score36

    Together AI's canary rollouts upgrade production models without downtime

    AITogether AI's canary rollouts shift production traffic between two model deployments on the same endpoint in staged percentages, with optional metric gates between steps. Operators can choose canary, blue-green, or rolling strategies, and a rollout starts only when explicitly launched; it can be paused, canceled, or reversed. The platform scales the target before moving traffic and waits for routing to converge before draining the source.

  4. Xiaomi MiMoOfficialAI score44

    MiMo-V2.6-Pro assists scientific research in materials and formal mathematics

    AIXiaomi's MiMo-V2.6-Pro, without research-specific RL training, helped Xiaomi materials researchers propose MOF materials for capturing PFAS "forever chemicals" and ran computational screening for wet-lab validation. It also helped formalize the full main theorem of Li–Yorke's "Period Three Implies Chaos" in Lean 4, producing a project of 6,000+ lines verified by Lean's kernel with no unfinished proof placeholders.

    Video from @XiaomiMiMo's post
  5. Gemini NotebookOfficialAI score26

    Gemini Notebook adds Interactive Learning Overviews for all users

    AIGemini Notebook now offers Interactive Learning Overviews to all users, according to the official account. The feature lets users combine source summaries with studio artifacts in a single interactive hub under Reports, aimed at studying and in-depth topic exploration.

    Video from @Gemini_Notebook's post
  6. Google AI DevelopersOfficialAI score22

    Google demos Gemini 3.8 coding tutor that sees your screen

    AIGoogle AI Developers showed a coding tutor built with Gemini 3.8 Live Extended Thinking that views the user's screen and calls functions to reference the p5.js library. The tutor points out the exact bug on screen and talks the user through the logic.

    Video from @googleaidevs's post
  7. Mike KnoopXAI score38

    Mike Knoop says LLM logprobs are vanishing, yet they enable useful new patterns

    AIMike Knoop notes that logprobs used to be widely exposed by LLM inference APIs and sees the market maturing so that parts of the LLM stack can be packaged in new, useful ways. He links this to Bryan Helmig's post on prompting with max_tokens: 1 plus logprobs for fast, parallel judgments, which Helmig says has a lot more depth than he expected.

  8. SemiAnalysisBlogAI score62

    How MoE inference splits into prefill, midfill, and decode regimes

    AIThe article explains how Mixture of Experts models change inference by making prefill, midfill, decode attention, and decode experts distinct workloads. It describes how KV cache state, expert routing, and parallelism choices shape compute, memory, and network demands across an inference cluster.

  9. LMSYS OrgOfficialAI score30

    LMSYS Publishes Blog Post on NVFP4 KV Cache Quantization

    AILMSYS Org shared a blog post about NVFP4 KV cache, a topic linked from its 2026-09-16 article. The post itself contains only a link, so no further technical details, figures, or results can be confirmed from this source.