Updated
#Tutorial/How-to
Updated
Sep 30
Hamel HusainAI score4 Hamel HusainAI score38 Hamel Husain Reviews Claude's New Auto Eval Plugin for Evaluations
AIHamel Husain has published a longer review of a new Claude Auto Eval plugin after many users asked about it. He invites readers to share their experiences using the plugin and how it went for them. The plugin is part of Claude's ability to help build evaluations and hillclimb on them, as described by @ClaudeDevs.
NVIDIA AIAI score12 NVIDIA to livestream building a visual AI agent from one prompt
AINVIDIA announced a livestream tomorrow at 9 a.m. PT in which it will build a visual AI agent from a single prompt. Viewers are invited to bring their questions to the session.
FireworksAI score34 GLM 5.3 Flash now available for training on Fireworks' Serverless API
AIFireworks AI has made GLM 5.3 Flash available for training through its Serverless Training API, open to all users. The model supports both vision and text inputs. Fireworks says it performs well on its benchmarks for agentic coding, document analysis, and tool use while remaining cost-efficient to serve.
Google Cloud TechAI score28 Agent Clinic Ep 3 builds automated eval suite for LangGraph agent
AITerminal test runs miss multi-turn agent regressions, so Agent Clinic Episode 3 builds an automated eval suite for a LangGraph agent in 60 minutes. The post presents a four-step framework for moving from informal checks to benchmarking AI agents, with a link to the full guide.
Google Cloud TechAI score4 Google Cloud Tech shares a guide to running open models cheaply in the cloud
AIGoogle Cloud Tech promotes a Medium article titled "How to run open models in the cloud without going broke." The post contains only a link and no further details about the methods, models, or costs discussed.
Google Cloud TechAI score15 Google Cloud tips for capping GPU and replica settings to control costs
AIGoogle Cloud recommends limiting accelerator count to a single GPU, setting replica count to 1-1, and avoiding capacity reservations to keep monthly bills predictable. These strict hardware limits apply to auto-scaling configurations for AI workloads.
Google Cloud TechAI score7 Google Cloud Spot VMs offer discounted compute for resilient developer and agent workloads
AIGoogle Cloud recommends Spot VMs for developer and agent testing workflows that tolerate interruptions and do not need continuous uptime. Spot VMs use spare Google Cloud compute capacity at deep discounts.
eric zakariassonAI score36 Eric Zakariasson shares a skill for building games with new models
AICursor's Eric Zakariasson shared a game-builder skill he has been using to build games with new models. The skill can be installed with the command npx skills add ericzakariasson/skills --skill game-builder.
NVIDIA AIAI score40 NVIDIA Shows Visual AI Agent Built in Under 30 Minutes
AINVIDIA says a single prompt can build and deploy a visual AI agent for a manufacturing line in under 30 minutes, with alerts, video search, and incident reports. The method uses the new Build Vision AI skill in NVIDIA VSS Blueprint 3.3, and a tutorial is available for readers who want to build one.
CognitionAI score16 Crosby Legal shares how it uses Devin, Cognition's AI agent
AICognition posted a link to a customer story about how Crosby Legal uses Devin, showing the approach the firm took. The post itself gives no further details on the results, workflow, or figures.
Ant LingAI score31 Ant Ling model turns plain-language prompts into interactive Three.js pages
AIAnt Ling can convert plain-language prompts into standalone, interactive Three.js pages covering topics such as an internal combustion engine, an optical-disc reader, paramecium organelles, and vector-field divergence. The output is runnable code rather than just an explanation.
Ant LingAI score38 Ling-3.1-flash ports C image library to Rust with 8.015× speedup
AIAnt Ling reports that its Ling-3.1-flash model completed a roughly 20-hour Rust port of a C image library. After a performance regression caused by busy-waiting workers and a parallelism adjustment, the model recovered and reached an 8.015× speedup. All 30 correctness checks passed.
NVIDIA AIAI score27 NVIDIA NeMo Relay Traces Hermes Agent Runs in Arize Phoenix
AINVIDIA and Nous Research published a hands-on walkthrough of NVIDIA NeMo Relay for collecting traces from Hermes Agent. The guide runs two example scenarios and shows the agent's calls and retries in Arize Phoenix. It also covers how Nous used traces and task results to evaluate fixes across repeated runs.
O'Reilly RadarAI score45 The Agentic Data Science Playbook: Delegating Analysis to AI Agents
AIAgentic data science has AI agents explore datasets, choose modeling approaches, run analyses, and explain findings while data scientists frame questions and verify evidence. In an experiment, Claude Opus 5.0 given the vague prompt "Build me a model to detect fraudulent nodes" on a modified Elliptic Bitcoin dataset reported F1 0.87 and ROC AUC 0.99 using a random split that leaked a planted label proxy.
Google GeminiAI score14 Gemini skills let you reuse instructions for presentations, emails, and brainstorming
AIGemini now lets users save custom skills to reuse instructions across tasks. The post suggests using them to outline presentation slides and prep Q&A, draft emails in a personal voice with Google Workspace context, and explore three to five perspectives on a topic.
Google GeminiAI score36 Gemini introduces skills, saved instruction sets for recurring tasks
AIGoogle Gemini now offers skills, which are saved sets of instructions for specific tasks you repeat often. The post describes them as shortcuts that let users skip repetitive prompting and go straight to results.
IdeogramAI score25 Ideogram 4.5 lets users restyle room photos with furniture edits
AIIdeogram 4.5 turns a single room photo into a design canvas where furnishings, finishes, and palettes can be swapped one edit at a time. The room's architecture stays unchanged throughout the edits.
IdeogramAI score20 Ideogram suggests multi-turn editing to restore old photos step by step
AIIdeogram recommends restoring an old photo in stages: first remove scratches and damage, then add color, then enhance. The post says trying to fix everything in a single prompt rarely works, so multi-turn editing lets each edit be refined separately.
Google Cloud · AI & Machine LearningAI score41 Google Cloud Rolls Out Agent Substrate, GKE Agent Sandbox RL Tools in September
AIGoogle Cloud introduced GKE Agent Substrate, an open-source execution runtime it says can run millions of sandboxes with 10x higher density than standard container runtimes. It also made GKE Agent Sandbox optimized for reinforcement learning generally available, alongside an orchestration SDK and native RL gym integrations. Google said GKE Pod snapshots can reduce AI inference start-up by as much as 89%, based on internal tests.
NVIDIA AIAI score20 NVIDIA Jetson introduces agentic development for building edge AI
AINVIDIA announced a new approach to building edge AI through agentic development on NVIDIA Jetson, presented in a broadcast. The post provides no further technical details, figures, or availability information.
Hamel HusainAI score15 Sampling production traces for review: explore broadly, then use signals
AIHamel Husain advises starting with broad exploration of production traces, then using signals to locate likely failures. He recommends keeping some random traces in every review batch so new problems can still be discovered.
KhazixAI score9 Blogger shares a checklist for keeping a new Claude account stable
AIThe author, whose earlier device was flagged so the account got banned within about half an hour, reports a new Claude account has run stably for six days. The shared tips include logging in with a Google account, using a home static IP, a clean new device, timezone set to Taiwan, paying via Google Play, starting at the $20 Max tier, and running Claude on a single always-on Mac Mini accessed remotely.
howie.seriousAI score22 Why local AI agents like Claude Code and Codex rely on shell access
AILocal and desktop agents such as Claude Code and Codex are powerful largely because they can use the shell, which connects them to the whole CLI ecosystem. The post lists tools including git, ffmpeg, curl, pandoc, gh, cron, and ssh as examples. It also says the video itself was produced by an agent operating the shell.
Karl's AI WattsAI score38 Can a menu bar tool switch models without losing session?
AIAfter switching models, can I continue the original session? If so, I really want to keep this menu bar, so I don't have to re-explain the project every time I change models.
Hamel HusainAI score42 Hamel Husain Tests Anthropic's Claude Eval Plugin on Leasing Assistant Traces
AIHamel Husain reviewed Anthropic's new build_eval and hill-climb commands in the claude-api plugin for Claude Code, finding it useful for discovering issues like human handoff, formatting, and voice agent problems. He criticized it for pushing evaluator creation before data review, asking for label validation in Markdown files, and bundling four failure checks into one broad call-transfer evaluator. Husain says he would hold off on using it for now.
Kling AI BlogAI score52 Kling 4.0 adds longer video, multimodal references, and a redesigned creation page
AIKling AI says Kling 4.0 will support native video generation from 3 to 30 seconds, with up to 10 keyframes for timing control. The full model is scheduled for October 2026, with early access to the new creation page rolling out from September 28.
Sep 29
GitHubAI score13 GitHub's most-starred repo's making explained in a new video
AIGitHub shared a video about the making of one of its most starred repositories. The post links to a YouTube video and gives no further details about the project or its development.
Google Developers BlogAI score47 Google Details Sparse Attention Speedup for Video Diffusion on TPUs
AIGoogle Developers Blog describes how Sparse VideoGen (SVG) routes video diffusion attention heads into spatial or temporal sparse masks and implements them as custom JAX and Pallas Splash Attention kernels on TPU v6e. In isolated single-chip tests with 75.6K tokens and 10 heads, the sparse variants retain about 38.87% of query-key pairs. The article argues that theoretical sparsity must be converted into hardware tile skipping to yield real speedups.
Julien ChaumondAI score25 Hugging Face blog shares a DGX Spark handbook guide from exolabs
AIJulien Chaumond, who owns Hugging Face, recommends a DGX Spark guide written by @0xSero and published on the Hugging Face blog. The post itself gives no further technical details about the handbook's contents.
Liquid AIAI score18 Liquid AI releases decision models with migration guide and demo tutorial
AILiquid AI has published documentation for its decision models, including a models page, a migration guide, and a road-decider demo tutorial on GitHub. The post contains links only and gives no model names, parameter counts, benchmark scores, or pricing.
Amjad MasadAI score34 How to design and build a harness for frontier performance at lower cost
AIAmjad Masad shared a post on designing and building a harness to reach frontier performance at a fraction of the cost. The post itself gives no specific models, benchmarks, or figures, and the linked article from @pirroh supplies the background context.
DeedyAI score42 Deedy shares a Claude Code workflow for AI video generation
AIDeedy describes a video generation pipeline built around Opus 5.5 in Claude Code, routing image, video, audio, and TTS models through OpenRouter's single API key. The workflow adds reference-image consistency, animatics before full renders, a critic skill that screenshots and transcribes output for QA, and ffmpeg for most editing.
DatabricksAI score22 Databricks rolls out frontier models to employees on Day 1 via Unity Gateway
AIDatabricks says it aims to give its employees the best models on launch day, quickly adopting new releases such as Opus 5.5 and GPT-6 Sol while tracking real-world usage and cost. Its AI engineering team uses Unity Gateway to manage access, spend, and model selection across thousands of employees, and to decide which models join its AI stack.
Sebastian RaschkaAI score22 Raschka's visual guide compares RNNs, CNNs, and transformers for text classification
AISebastian Raschka has published a comprehensive write-up on language models for text classification, covering RNNs, CNNs, transformers, and calibration. It is presented as a visual guide with hands-on experiments on accuracy and efficiency.
howie.seriousAI score32 Wording-level prompt tricks are obsolete in 2026, author argues
AIThe author argues that carefully crafted wording-level prompts have almost no effect in 2026, and that clear intent plus sufficient context matters most. Reusable prompt components are being absorbed into agent skills and context tools, while harnesses and models internalize more capability, leaving little room for prompting.
howie.seriousAI score18 100-Second Git Explainer for the Agent Era
AIQuick video: "Git in 100 Seconds: What Everyone Should Know in the Agent Era" === Though most followers already know git, I made the video anyway, so here it is 🤣
Ahead of AI (Sebastian Raschka)AI score43 Language Models for Text Classification: From Bag-of-Words to Jev
AISebastian Raschka traces text classification from bag-of-words models such as naive Bayes and logistic regression through pre-transformer neural networks, then sets up an analysis of the recently released Jev AI model. The article frames Jev as a general-purpose classifier that trades specialized accuracy for speed, cost, and breadth of tasks.
IEEE Spectrum · AIAI score14 IC-STAR Brings Full-Flow Autonomous AI to Digital and Analog Chip Design
AIThe webinar presents IC-STAR, an autonomous AI approach that shifts silicon engineers from manually managing tools and handoffs to defining objectives and supervising AI-driven execution across the chip development lifecycle. It covers four enabling technologies and includes a look at Ambiq's production deployment of autonomous AI. The source provides no performance figures or availability details.
Thomas WolfAI score29 Thomas Wolf calls a post simply "impressive"
AIThomas Wolf, owner of the Hugging Face account, posted the single word "impressive" in response to a quoted post. The quoted post reports a new NanoGPT training record of 39.9s, down 27.7s from the prior 67.6s, achieved through per-flop optimizations such as sampled softmax and sparse updates.