Skip to contentSkip to stories

Updated

#Tutorial/How-to

Showing low-relevance items too. Hide low-relevance items

Sep 22

Sep 22Tue
  1. Together AI BlogOfficialAI score38

    How to train your own Jev classifier for $17 with Together AI

    AIThe Together AI blog shows how to fine-tune a Qwen3.5 4B base model into a classification model using about 38,000 examples sampled from six Hugging Face datasets, at a training cost of roughly $17.0. The tutorial covers cloning the tev1 repository, normalizing data with provided scripts, launching a Together AI fine-tuning job that takes about 25 minutes, and deploying the result to a dedicated H100 endpoint.

  2. Boris ChernyXAI score42

    Boris Cherny Uses Opus 5.5 to Formally Verify Claude Agent SDK

    AIBoris Cherny used Opus 5.5 to formally verify the Claude Agent SDK with Lean, and short prompts produced 16 PRs fixing bugs and race conditions. He also combines Lean and TLA+ to find issues in data flow, concurrency, and state management, and says Claude is strong in both languages even though he does not know them well.

    Video from @bcherny's post
  3. Alex AlbertXAI score37

    Claude prompt recreates 1906 Market Street in Blender for video

    AIA prompt shared by Alex Albert asks Claude to recreate San Francisco's Market Street as it stood on April 17, 1906, before the earthquake, using Blender. It requires building a source file from Sanborn fire insurance maps, the Miles Brothers film, period photos, and USGS topography, with reusable Blender Python generators for facades, street lamps, and vehicles, ending in a 10-second video up the street.

  4. Alex AlbertXAI score18

    Opus 5.5 builds a 1906 San Francisco street scene in Blender

    AIAlex Albert says Opus 5.5 has improved 3D modeling and vision for Blender work, letting users build an entire world from a single prompt. He shares a historically accurate render of San Francisco's Market Street in 1906, before the earthquake.

    Video from @alexalbert__'s post
  5. StepFunOfficialAI score31

    Step Code installs on macOS, Linux, and WSL via one command

    AIStepFun released an install script for Step Code that runs on macOS, Linux, or WSL with a single curl command. WSL is the recommended route on Windows, while PowerShell support is currently in beta. The project accepts issues and pull requests on GitHub.

  6. Unsloth AIOfficialAI score70

    Qwen-Image-2.1 runs locally on 12GB VRAM using Unsloth GGUFs

    AIUnsloth says the 7B Qwen-Image-2.1 text-to-image and editing model can run locally on 12GB VRAM using its GGUF builds. It also states that the model performs on par with Nano Banana 2.0, and that Dynamic FP8 can run on 6GB of VRAM via offloading for higher quality. The image lists int8 at 7.26 GB with mean LPIPS 0.064 and fp8 at 7.12 GB with mean LPIPS 0.112, and says int8 is the default.

    Why it matters: The post gives concrete local-run settings, VRAM figures, and GGUF and FP8 options, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  7. Kimi.aiOfficialAI score46

    Kimi launches browser extension for chatting, automating web tasks

    AIKimi has released its Kimi Browser Extension, formerly Kimi WebBridge, which runs in the browser sidebar to navigate websites and fill out forms. Users can record repetitive steps once and save them as a skill for Kimi to reuse later. The extension is available now on the Chrome Web Store.

    Video from @Kimi_Moonshot's post
  8. OpenBMBOfficialAI score20

    OpenBMB praises MiniCPM5-2B workers in multi-agent invoice reconciliation

    AIOpenBMB thanked a developer for testing MiniCPM5-2B as a worker in a multi-agent workflow handling invoice matching, short payments, duplicate references, and disputes through tool calls. The background post says GPT-6 Astra coordinated the MiniCPM5-2B workers, verifying 32 synthetic invoices in 67.8 seconds with 232 executed tool calls. The demo does not move money.

  9. Suno BlogOfficialAI score16

    Suno Studio's Compressor Reduces Dynamic Range to Balance Mixes

    AISuno Studio includes a compressor plugin that reduces dynamic range by turning down peaks and applying makeup gain to raise quieter parts. Its controls include threshold, ratio, attack, and release, and it is especially useful for shaping uploaded or recorded audio. Studio is browser-based and comes free with every Suno Premier subscription.

Sep 21

Sep 21Mon
  1. Tencent HyOfficialAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

  2. xAI News (Grok)OfficialAI score46

    How SpaceXAI uses Grok Bot to scale customer support without new hires

    AISpaceXAI says its combined support team handled a 175% rise in tickets without hiring, crediting Grok Bot, which it says would otherwise have required about 200 additional staff. The company reports resolving tickets for $0.20 to $0.30 each, versus the $1 to $4 per resolution it attributes to traditional AI support tools. Grok Bot is also reported to resolve 99% of refund requests without human intervention.

  3. Together AI BlogOfficialAI score36

    Together AI's canary rollouts upgrade production models without downtime

    AITogether AI's canary rollouts shift production traffic between two model deployments on the same endpoint in staged percentages, with optional metric gates between steps. Operators can choose canary, blue-green, or rolling strategies, and a rollout starts only when explicitly launched; it can be paused, canceled, or reversed. The platform scales the target before moving traffic and waits for routing to converge before draining the source.

  4. Xiaomi MiMoOfficialAI score44

    MiMo-V2.6-Pro assists scientific research in materials and formal mathematics

    AIXiaomi's MiMo-V2.6-Pro, without research-specific RL training, helped Xiaomi materials researchers propose MOF materials for capturing PFAS "forever chemicals" and ran computational screening for wet-lab validation. It also helped formalize the full main theorem of Li–Yorke's "Period Three Implies Chaos" in Lean 4, producing a project of 6,000+ lines verified by Lean's kernel with no unfinished proof placeholders.

    Video from @XiaomiMiMo's post
  5. Gemini NotebookOfficialAI score26

    Gemini Notebook adds Interactive Learning Overviews for all users

    AIGemini Notebook now offers Interactive Learning Overviews to all users, according to the official account. The feature lets users combine source summaries with studio artifacts in a single interactive hub under Reports, aimed at studying and in-depth topic exploration.

    Video from @Gemini_Notebook's post
  6. Google AI DevelopersOfficialAI score22

    Google demos Gemini 3.8 coding tutor that sees your screen

    AIGoogle AI Developers showed a coding tutor built with Gemini 3.8 Live Extended Thinking that views the user's screen and calls functions to reference the p5.js library. The tutor points out the exact bug on screen and talks the user through the logic.

    Video from @googleaidevs's post
  7. Mike KnoopXAI score38

    Mike Knoop says LLM logprobs are vanishing, yet they enable useful new patterns

    AIMike Knoop notes that logprobs used to be widely exposed by LLM inference APIs and sees the market maturing so that parts of the LLM stack can be packaged in new, useful ways. He links this to Bryan Helmig's post on prompting with max_tokens: 1 plus logprobs for fast, parallel judgments, which Helmig says has a lot more depth than he expected.

  8. SemiAnalysisBlogAI score62

    How MoE inference splits into prefill, midfill, and decode regimes

    AIThe article explains how Mixture of Experts models change inference by making prefill, midfill, decode attention, and decode experts distinct workloads. It describes how KV cache state, expert routing, and parallelism choices shape compute, memory, and network demands across an inference cluster.

  9. LMSYS OrgOfficialAI score30

    LMSYS Publishes Blog Post on NVFP4 KV Cache Quantization

    AILMSYS Org shared a blog post about NVFP4 KV cache, a topic linked from its 2026-09-16 article. The post itself contains only a link, so no further technical details, figures, or results can be confirmed from this source.

  10. Matei ZahariaXAI score32

    Matei Zaharia praises GEPA working with Jev

    AIMatei Zaharia, a prominent AI researcher, said it is very cool that GEPA works on Jev. The post is a short endorsement, linking to background about a test in which GEPA optimized Jev's prompts for extracting suspected adverse drug effects from medical sentences.

  11. Ali GhodsiXAI score14

    Superhuman scales its GEC inference infrastructure, per Ali Ghodsi

    AIDatabricks CEO Ali Ghodsi praises how Superhuman scales its AI infrastructure, linking to a Superhuman blog post on scaling GEC inference. The post itself is only a link with no details provided here, so the specific techniques and results cannot be confirmed from this source.

Sep 20

Sep 20Sun
  1. vLLM BlogOfficialAI score44

    vLLM Reports PD Serving Results for Qwen3.8-2.4T on GB300 NVL72

    AIvLLM achieved 5000 total token throughput per GPU in high-throughput PD serving of Qwen3.8-2.4T on a GB300 NVL72 cluster under an 8K/1K workload. The low-latency scenario reached 180 generated tokens per user, with both results shown on the Pareto frontier. The post also provides srt-slurm recipes and explains the tuning process used to create them.

  2. WanOfficialAI score12

    Wan Lab offers a new skill for trying Wan video creation

    AIWan has released a skill on its create.wan.video lab platform that users can try now. The post links directly to the skill page and gives no further details on its features, capabilities, or pricing.

    Image from @Alibaba_Wan's post

Sep 19

Sep 19Sat
  1. OpenBMBOfficialAI score34

    OpenBMB's 2B MiniCPM5 powers a local personal news desk

    AIOpenBMB's 2B-parameter MiniCPM5 model runs as a local news desk on an older i5-9400F PC with 16GB RAM and no cloud API. The developer built a system that collects official sources hourly and sends a 24-hour Telegram recap with a lead story and links.

  2. Sebastian RaschkaXAI score36

    Raschka's Inference Scaling Part 1: Sampling for Better Accuracy

    AISebastian Raschka starts a series on inference scaling by modifying text generation with temperature scaling, top-p filtering, and multinomial sampling to produce diverse outputs. He says this enables self-consistency and best-of-N approaches that improve answer accuracy by more than 2x. The video covers chain-of-thought prompting, a MATH-500 evaluation, and accuracy versus compute tradeoffs.

    Video from @rasbt's post

Sep 18

Sep 18Fri
  1. TinkerOfficialAI score31

    Jasper's guide shows how reward tweaks shape search agent behavior

    AIJasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency. The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.

  2. Google AI DevelopersOfficialAI score22

    Gemini 3.8 Live Extended Thinking guides first-time cyberdeck builders

    AIGoogle AI Developers demonstrates Gemini 3.8 Live Extended Thinking as a workshop assistant that analyzes a user's workspace, reasons aloud, and guides a beginner through building a cyberdeck. The demo targets users with no prior experience building one.

    Video from @googleaidevs's post
  3. LMSYS OrgOfficialAI score16

    LMSYS releases SGLang SSD expert pack blog post

    AILMSYS Org published a blog post introducing an SGLang SSD expert pack, with the full details available on its website. The post itself gives no further technical specifics, so the summary is limited to the announcement.

  4. LlamaIndex 🦙OfficialAI score16

    LlamaParse Preserves Table Structure in EIA Energy Report Data

    AILlamaIndex launches a Parsed by LlamaParse series, using the EIA's September 2026 Short-Term Energy Outlook to show how table parsing errors can corrupt downstream data. The example value 1,186 in Table 7a means electricity sales to ultimate customers in Q3 2026, in billion kilowatthours, and misparsing its quarter, metric, or unit could flow into dashboards and forecasts. The post says LlamaParse preserves structure and footnote context needed for databases, forecasting, and AI applications.

    Image from @llama_index's post
  5. Google · AI blogOfficialAI score29

    Google co-builds Google Flow tools with two designers for New York Fashion Week runways

    AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.

  6. MiniMax Design (H3)OfficialAI score18

    Hailuo AI speeds up storyboarding with 3x3 panel-to-video generation

    AIHailuo AI's post suggests that a single 3x3 storyboard image can be converted into a video using the MiniMax H3 Max r2v model at 480p for 15 seconds. The quoted post describes settings with Quality prompt tuning and standard reference strength, and asks for a 2D animation with panel-to-panel cuts while excluding multiple panels and BGM.

  7. Hamel HusainBlogAI score62

    Hamel Husain's FAQ on AI evals: error analysis, judges, and trace review

    AIHamel Husain and Shreya Shankar's FAQ explains AI evals as tests of whether an AI system does what users and the business want. It recommends starting with error analysis on at least 30 traces, then turning recurring failures into binary code-based checks or LLM judges validated against human labels.