Skip to contentSkip to stories

Updated

#Tutorial/How-to

Showing low-relevance items too. Hide low-relevance items

Sep 13

Sep 13Sun
  1. Sebastian RaschkaAI score35

    Raschka's Reasoning from Scratch Round 3 Builds a Math Verifier

    AISebastian Raschka's third "Reasoning from Scratch" video covers building a math verifier for evaluating language models and for later reinforcement learning with verifiable rewards (RLVR) training. The walkthrough covers extracting final answers from boxed outputs, normalizing them, checking mathematical equivalence, and running evaluation on the MATH-500 dataset.

    Video from @rasbt's post

Sep 11

Sep 11Fri
  1. Augment Code BlogAI score80

    Augment Code details how its software factory raised output per developer 4.5×

    AIAugment Code reports that size-adjusted output per active developer rose from 12.3 to 55.7 between November 2025 and July 2026, while median time to merge fell from 11.2 to 3.1 hours. The post says the company added specialized agents wherever work was piling up, across planning, review, verification, feedback, and incident response, and kept engineers responsible for product decisions, architecture, and production risk.

    Why it matters: The post pairs internal productivity and quality metrics with the order in which agents were added, showing how review and verification bottlenecks shaped a software delivery pipeline.

Sep 10

Sep 10Thu
  1. Google Developers BlogAI score55

    Google details autonomous LLM post-training loops using Tunix on TPUs

    AIGoogle Developers Blog describes autofinetune, a project applying autonomous agent loops to LLM post-training with Tunix, Gemma, and Cloud TPUs. In an SFT case study on FunctionGemma, an agent ran 20 automated experiments on a Cloud TPU v5e-1 to adjust LoRA settings, optimizers, and learning rates. In a GRPO case study on Gemma 3 1B for GSM8K math reasoning, the agent ran 40 experiments on a Cloud TPU v6e-1 and improved total reward by about 10%.

Sep 9

Sep 9Wed
  1. Fireworks AI BlogAI score58

    Fireworks AI outlines a staged path from closed APIs to owned specialized models

    AIFireworks AI describes a four-stage path for teams moving from renting closed frontier models to training their own, starting with API use and prompt, context, and harness engineering. The post uses the UIPad computer-use dataset to show that Kimi K3 ties GPT 5.6 Sol overall at 87.7 but wins three of four categories while costing about half as much, suggesting routing. After roughly three hours of training on the training split, the tuned Kimi K3 outperforms GPT 5.6 Sol on the held-out test set.

  2. Mistral AIAI score54

    Mistral details how AI agents migrated 40,000 lines of Fortran to C++

    AIMistral AI helped a European energy operator migrate 40,000 lines of Fortran 77 to C++ for a reservoir simulator with no test suite. The post explains a parity harness that checks numerical agreement between the two codebases, and a workflow where agents coder, tester, and reviewer migrate modules under human review. Its authors note the approach covered the self-contained first sprint of 40,000 of 300,000 lines and that dependent systems would bring additional challenges.

Sep 8

Sep 8Tue
  1. Google Developers BlogAI score36

    Google Developers Blog outlines behavioral evals for guarding AI coding agents against regressions

    AIGoogle Developers Blog argues that teams building AI coding agents should replace end-to-end benchmark scores with behavioral evaluations that test discrete, observable actions. Examples include asking clarifying questions on underspecified prompts, running a local validator before marking a build change complete, and consulting live search for current information. The post recommends fast, deterministic unit-style checks, outcome-based LLM-as-a-judge checks for complex tasks, and batch runs that track aggregate pass rates over time.

  2. InferactAI score42

    Inferact reports open models hit 130K tokens/GPU-sec on agentic workloads

    AIInferact says months of vLLM tuning for agentic workloads, validated on SemiAnalysis's AgentX benchmark, let open-source models reach up to 130K tokens per GPU-second. The company claims this is 106 times cheaper than Opus 5 API pricing. The work is described as part of a vLLM blog post covering architecture, framework, and runtime optimizations.

Sep 6

Sep 6Sun
  1. Sebastian RaschkaAI score22

    Raschka's Reasoning From Scratch video covers LLM text generation and KV caching

    AISebastian Raschka released a video in his Reasoning From Scratch series covering text generation in LLMs and KV caching. The walkthrough uses a pretrained Qwen3 model from the Reasoning From Scratch package, covering tokenization, greedy decoding, end-of-sequence handling, and a benchmarked KV caching speedup. It prepares the base model for reasoning techniques in later episodes.

    Video from @rasbt's post

Sep 4

Sep 4Fri
  1. Andrew NgAI score42

    Andrew Ng maps key skills for using AI coding agents effectively

    AIAndrew Ng presented an AI Engineering Skills Map for using coding agents such as Claude Code, Codex, Cursor, OpenCode, and Pi. The workflow he describes covers planning, execution, and deployment with monitoring, and he identifies five key skills: directing the workflow, enabling agent autonomy, reviewing the work, customizing the agent and its environment, and coding agent foundations. The source says these skills matter more as the agents evolve quickly.

Sep 3

Sep 3Thu
  1. TinkerAI score23

    Tinker used to test counterfactual simulatability for LLM interpretability

    AITinker, the platform from @tinkerapi, supported two recent papers testing counterfactual simulatability as a way to interpret LLM behavior. The core idea is that understanding a model means predicting how its output changes when the prompt changes, with causes ranging from specific words to abstract properties such as a user's angry tone.

  2. Google Developers BlogAI score23

    Google's Gemini Enterprise DevEx sprint fixes governance setup friction for agents

    AIGoogle's Gemini Enterprise developer experience team tested agent governance workflows without internal shortcuts and fixed friction points across its agent governance products. Fixes included documentation stating that enabling the Identity-Aware Proxy API is a hard requirement, auto-allowing essential Google-managed platform APIs in the Agent Gateway, and adding Private Service Connect and Cloud DNS setup guidance for Semantic Governance. The team also published ready-made Logs Explorer queries for monitoring Agent Gateways and Content Security.

  3. Matei ZahariaAI score34

    Databricks uses Unity AI Gateway traces to cut AI waste fast

    AIDatabricks used Unity AI Gateway tracing and Genie One to find seven small MCP-server bugs and eliminate an estimated $1.2M in annual wasted AI spend and lost productivity within an hour. The bugs drove about $499K per year in wasted tokens, roughly 12,000 engineering hours per year in agent wait time, and 1,409 tool errors in a single 24-hour window. Matei Zaharia argues that analyzing tracing data for AI workloads will become a routine form of operational data analysis across companies, much like finance and security.

  4. Prime Intellect BlogAI score59

    Prime Intellect rebuilds GLM-5.2 RL weight transfer on NIXL, cutting sync to 3.9 seconds

    AIPrime Intellect reports that rebuilding RL weight transfer for GLM-5.2 on NIXL and ModelExpress cut sync time from 86.1 seconds with NCCL to 3.9 seconds in its fastest setting. The method traces vLLM's loader to find each tensor's runtime layout, then reads only the needed source bytes over RDMA and replays the rest locally. Most remaining latency comes from vLLM's pause consensus, which the team reduced by syncing every wave instead of every 32.

Sep 2

Sep 2Wed
  1. Daniel HanAI score34

    Stanford's Modern Software Developer course adds AI-native engineering curriculum

    AIMihail Eric announced the 2026 edition of his Stanford course "The Modern Software Developer," with 85% of the Fall 2025 material replaced by AI-native topics such as agent skills, context engineering, and agentic code review. Students will ship pull requests to real open-source AI repositories, with partners including Browserbase, HeyGen, and CopilotKit offering mentorship.

  2. Engineering at MetaAI score55

    Meta details an AI agent that learns from expert corrections without retraining

    AIMeta Engineering describes an AI agent for a compliance domain that stores expert knowledge in structured, auditable files and separates it from reasoning procedures called recipes. Expert feedback is diagnosed, compiled into verified text edits, tested against regression suites, and reviewed by humans, all without retraining the underlying model. Meta reports that domain experts rated outputs useful almost all the time and that assessment time fell from days to minutes.

Sep 1

Sep 1Tue
  1. Google Developers BlogAI score39

    Four engineering patterns behind top Google AI Agents Challenge submissions

    AIGoogle's AI Agents Challenge judges highlighted four engineering patterns in top-ranked submissions: bidirectional MCP, event-driven concurrency, same-bar fallback, and tiered routing. One team exposed its internal MCP tools as an external MCP server that other agents could call, with access control required once outside callers reach it. Another replaced a linear agent pipeline with an asyncio.Queue-based event bus so agents react to shared events in parallel rather than waiting in a call chain.