Skip to contentSkip to stories

Updated

#Tutorial/How-to

Aug 13

Aug 13Thu
  1. Augment Code BlogAI score22

    Augment Code uses Cosmos to check enterprise pilot health against usage and deal data

    AIAugment Code's Solutions Architecture lead used the Cosmos agentic orchestration platform to build a live pilot-health view that combines product usage, GitHub and PR activity, Salesforce deal data, and customer call transcripts. Each account's health and board-level one-liner was checked against the customer's own stated success criteria, such as a 30% PR merge-time reduction. The article says the view refreshed from current Salesforce data and was designed to avoid inflating usage numbers through session lineage reconciliation.

Aug 11

Aug 11Tue

Aug 10

Aug 10Mon

Aug 7

Aug 7Fri
  1. Matei ZahariaAI score44

    AI tokens are another resource to optimize in software engineering now.

    AIMy cofounder @pwendell wrote about how we and other tech companies are starting to manage this resource now, by routing everything through an AI Gateway. This enables centralized analysis (e.g. we found settings we can change on Claude Code and Codex to lower cost a lot), smart routing, and “pushing down” control to our engineers so they can set budgets on individual tasks and prevent surprises.

  2. Ali GhodsiAI score58

    Databricks details four techniques it used to cut internal AI coding spend by up to 90%

    AIDatabricks published an analysis of four techniques it used to reduce internal AI spend while growing adoption, with savings of up to 90% in some scenarios. The techniques are shifting defaults to cheaper models such as GLM, automated task-level model routing, per-user spend visibility with adaptive budgeting, and pruning context bloat. The author, Ali Ghodsi, reposted Databricks co-founder Patrick Wendell's summary and recommended it.

Aug 5

Aug 5Wed

Aug 4

Aug 4Tue

Aug 3

Aug 3Mon

Aug 2

Aug 2Sun
  1. OpenRouter BlogAI score40

    OpenRouter Launches Ori Eval to Find the Best AI Model for Your App

    AIOpenRouter has released Ori Eval, an agent-driven tool that runs your app's prompts against candidate models and returns a comparison table of catch rate, latency, cost per PR, and pass/fail results. The tool asserts on called tools and grades open-ended answers with an LLM judge, pinning the harness and model during each run. Its evals are code files that can run in CI to block regressions and re-run when new models ship.

Jul 29

Jul 29Wed
  1. Fireworks AI BlogAI score54

    Fireworks tests whether LoRA or full fine-tuning gaps come from data, learning rate, or rank

    AIFireworks AI ran controlled SFT experiments on Qwen3.5-9B comparing LoRA with full parameter fine-tuning across three synthetic verifiable tasks. The post argues that a FullFT advantage can come from data coverage, learning-rate tuning, or adapter rank, and it recommends testing these in that order before switching methods. Under a fixed multi-task budget, FullFT kept a 4.29-point lead over the best LoRA recipe tested, while matched data exposure favored LoRA.

  2. Liquid AI NewsletterAI score46

    Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi

    AILiquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens. The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages. LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.

Jul 28

Jul 28Tue
  1. Fireworks AI BlogAI score46

    Fireworks AI Shows Low-Cost Fine-Tuning Lifts Domain Embedding Retrieval

    AIFireworks AI describes fine-tuning Qwen3-Embedding-8B on private (query, positive) pairs using bidirectional InfoNCE loss through its Training SDK, then serving the model via an OpenAI-compatible embeddings endpoint. The post reports that around 150 training steps was enough, that rank-32 LoRA landed within about one point of full-parameter fine-tuning, and that gains were largest where the base model struggled, while tasks like CoSQA and FiQA2018 showed flat results.

Jul 25

Jul 25Sat

Jul 24

Jul 24Fri

Jul 23

Jul 23Thu
  1. One Useful Thing (Ethan Mollick)AI score67

    Ethan Mollick's guide to choosing AI tools for agentic work

    AIEthan Mollick's guide says ChatGPT and Claude are the main choices for real work, since their agent modes can act on a computer. He separates agent modes that run on the company's computers from those that access the user's own computer. He recommends keeping approval settings on for sending, spending, or deleting, because of prompt injection risk. He also notes that Gemini currently lags for agentic work, though its Notebook and video tools are useful.

Jul 22

Jul 22Wed

Jul 21

Jul 21Tue
  1. OpenAI NewsroomAI score14

    With nine kids and a no-screens rule on Sundays, Ben Kalkman wanted a project his whole family could build together: a new treehouse.

    AIBen used ChatGPT as his construction consultant, helping Ben shape the structural design, source materials, and, crucially, translate his kids’ “wild ideas” into actual, buildable plans. Today, the treehouse is still evolving, and the family project keeps growing one Sunday - and one wild idea - at a time.

  2. Andrej KarpathyAI score30

    Karpathy suggests long voice rambles help LLMs understand your intent

    AIAndrej Karpathy describes using /voice to ramble for about 10 minutes, sometimes as a short interview, to give an LLM context that would be tedious to type. He says LLMs reconstruct these messy streams of thought remarkably well, often returning a cleaner version than the speaker started with, which improves shared understanding and reduces later corrections.

Jul 19

Jul 19Sun

Jul 18

Jul 18Sat
  1. Ahead of AI (Sebastian Raschka)AI score52

    How Reasoning Effort Settings Are Built Into LLMs Through Training

    AIThe article explains how reasoning models can offer multiple effort modes, separating training-time methods from inference-time controls such as system prompts and chat templates. It compares six open-weight models, including DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5, GLM-5, Qwen3, and Inkling, noting that their reports disclose different levels of detail. It also shows how GPT-5.6's model selection and effort settings act as two separate scaling axes.

Jul 17

Jul 17Fri
  1. OpenAI NewsroomAI score22

    Foreguard, built with ChatGPT and Codex, helps families plan care and benefits early

    AISekhar and Katie Brandt built Foreguard, a free tool built with ChatGPT and Codex, to help families claim public benefits they are entitled to and plan private insurance coverage. The tool shows that modest budgets of $100 a month can create millions of dollars of day-one financial protection. The goal is to help families prepare earlier, before care decisions become urgent.

  2. Andrew NgAI score28

    DeepLearning.AI launches course on fast LLM inference with Cerebras

    AIDeepLearning.AI has launched a short course, built with Cerebras, on building LLM applications that respond quickly using inference-optimized hardware. The course compares how GPUs, TPUs, and Cerebras' Wafer-Scale Engine handle the memory-to-compute bottleneck, which keeps model weights close to compute units to speed token generation. It covers real-time applications such as live translation and voice agents, plus habits for agentic coding.

Jul 16

Jul 16Thu

Jul 14

Jul 14Tue

Jul 12

Jul 12Sun
  1. OpenAI NewsroomAI score12

    James Costello uses ChatGPT to run his demolition business

    AIStructural engineer James Costello, who oversees complex New York City high-rise demolitions, uses ChatGPT to review lengthy contracts, organize compliance documents, and create construction plans for his family-rooted firm DEMTEC. The post says the tool helps him move through these workflows faster and with more confidence, freeing time to grow the business and support his team.

  2. OpenAI NewsroomAI score8

    Big Chan was making moves in hip-hop, but her biggest breakthrough came after the music stopped.

    AIWith her career at a crossroads, she moved to Atlanta, found a sewing machine in a friend’s attic, and taught herself a new craft. Inspired to design by her mother’s legendary disco style, Chan uses ChatGPT to help refine her sketches and bring her visions to life. She continues to create her own opportunities within the fashion industry, now dressing some of entertainment’s biggest names.

Jul 11

Jul 11Sat
  1. OpenAI NewsroomAI score15

    Emma Dahl used ChatGPT to help design and build her custom wedding dress

    AIEmma Dahl wanted a historically inspired wedding dress that incorporated pearls from her grandmother's necklace, so she used ChatGPT over months to troubleshoot niche sewing and corset construction. The chatbot also helped her choose a sewing machine upgrade, the shape of her veil, and how to pack and travel with the orchids for her bouquet.

Jul 8

Jul 8Wed

Jul 6

Jul 6Mon

Jul 5

Jul 5Sun
  1. Cat WuAI score18

    what are your top claude code + workflows + artifacts use cases?

    AIone of my new favorites is for sourcing candidates: - tell cc about the role and backgrounds i’m looking for - ask cc to kick off a dynamic workflow to find 100 candidates, and include linkedin, twitter, blog, podcasts, and one-line pitch for each - ask cc to make an artifact and email it to me. then, i lock my laptop and head out for the day - i review the list on the go once cc is done!

Jun 30

Jun 30Tue
  1. Andrew NgAI score50

    Andrew Ng outlines three loops for building 0-to-1 AI products

    AIAndrew Ng describes three loops he uses to build 0-to-1 products with AI agents: an agentic coding loop, a developer feedback loop, and an external feedback loop. He says the agentic coding loop runs every few minutes, letting coding agents build, test, and iterate on software for around an hour without human intervention. The developer feedback loop operates over tens of minutes to hours, with humans steering product decisions because they hold a context advantage over AI systems.

Jun 27

Jun 27Sat
  1. Ahead of AI (Sebastian Raschka)AI score37

    Local Coding Agents: Setting Up Qwen3.6 with Open-Source Harnesses

    AISebastian Raschka's tutorial shows how to build a fully local coding agent by pairing an open-weight LLM served through an inference runtime with an open-source harness that can read files, edit code, and run commands. He recommends Qwen-Code for Qwen3.6, citing Nvidia's Polar paper, which found Qwen models performed best in Qwen-Code. The Qwen3.6 35B-A3B model is about 22 GB to download and needs roughly 30–40 GB of RAM.