Skip to contentSkip to stories

Updated

#Deployment/Engineering

Showing low-relevance items too. Hide low-relevance items

Sep 24

Sep 24Thu
  1. Baseten BlogOfficialAI score44

    LangSmith Fine-Tuning Trains Open Models on Agent Traces via Baseten Loops

    AILangChain launched LangSmith Fine-Tuning, which lets users fine-tune open models on their LangSmith agent traces using the open-source smithtune CLI. Training runs on Baseten Loops in the user's own workspace, and smithtune deploy places the evaluated checkpoint on a Baseten Dedicated Inference deployment. Loops is in early access, so users may need to request access for their workspace.

  2. Azure BlogOfficialAI score67

    Microsoft Foundry adds voice agents and continuous optimization for production agents

    AIMicrosoft Foundry expands its agent platform with voice agents in public preview, long-running resilience for hosted agents, and tools for evaluating production agents. The post also says GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are now available in Foundry. Agent optimizer, Insights, and Rubric evaluator are described as tools for continuous improvement, with some reaching general availability later this month.

    Why it matters: The post shows how Foundry combines model choice, voice agents, long-running resilience, and production evaluation into one agent workflow, with a customer example.

  3. Microsoft Foundry BlogOfficialAI score40

    Foundry Agent Service adds egress policies to restrict hosted agent destinations in preview

    AIMicrosoft's Foundry Agent Service preview lets developers attach a named, ordered egress policy to a hosted agent, allowing only approved destination hostnames. The walkthrough uses an invoice agent, an Audit-mode RAI policy with a Deny default, and Allow rules for two finance and vendor hosts, configured outside the agent code. Network egress controls are preview features, not GA, with no preview SLA, and are not intended for production use.

  4. Google ResearchOfficialAI score38

    Google's John Platt on AI for climate, disease forecasting, and science

    AIIn a Latent Space podcast episode, Google's John Platt discusses using AI to address climate change, including reducing airplane contrails that contribute about 1% of human-caused warming and detecting fires with FireSat satellites. He also describes Google's Empirical Research Assistance (ERA), which uses Gemini and Monte Carlo Tree Search and achieved top marks in recent CDC benchmarks for forecasting COVID and flu cases a week ahead.

  5. LiveKitOfficialAI score28

    LiveKit tests Gemini 3.8 Flash-Lite TTS in a live voice agent

    AILiveKit tested Gemini 3.8 Flash-Lite TTS inside a LiveKit agent, letting users direct a voice line by line and hear it hold up in a real conversation. The post highlights expressive speech, custom voices, and a production-ready voice library.

  6. Epoch AI · The Epoch BriefOfficialAI score45

    Huawei Trails Nvidia by About Four Years in AI Chip Performance and Output

    AIHuawei will likely remain about four years behind Nvidia in AI chip performance and production through 2030, Epoch AI estimates. Its flagship Ascend 950 delivers roughly half the performance of Nvidia's 2022 H100, and Huawei is projected to produce about 1.5 million chips in 2026 versus Nvidia's roughly 6 million, leaving it about 25 times behind in total compute.

  7. Google · Gemini appOfficialAI score62

    Google launches Gemini 3.8 Live with Live Avatar for enterprises

    AIGoogle introduced Gemini 3.8 Live with Live Avatar, which adds a visual persona with lip-syncing and expressions to its live dialogue models. The feature is available in Gemini Enterprise and supports 97 languages, with custom avatars available through enterprise allowlisting. Google says all output is watermarked with SynthID.

    Why it matters: The post specifies enterprise availability, custom avatar allowlisting, and 97-language support, which clarifies who can use the feature and how far it reaches.

  8. vLLMOfficialAI score34

    vLLM and RL-Kernel achieve bit-exact logprob match on AMD MI300X

    AIThe RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime, and a 200-step Qwen3-8B GRPO run on 8× AMD MI300X recorded zero logprob mismatches between Megatron training and vLLM rollout. The strict path aligns reduction order, intermediate precision, rounding points, and math primitives across both sides to achieve bit-for-bit matching on ROCm.

  9. Microsoft Foundry BlogOfficialAI score61

    Microsoft Foundry Routines reach general availability for scheduled and event-driven agents

    AIMicrosoft announced general availability of Routines in Foundry Agent Service, a managed way to run agents on a timer, on a recurring schedule, or in response to GitHub issue events and new Microsoft Teams channel messages. Routines keep the trigger, agent action, identity, connections, and run history in the Foundry project, and each routine can run under the creator's identity or the agent's own Microsoft Entra ID identity. A preview reminder tool lets a Hosted Agent schedule itself to resume later on the same conversation.

    Why it matters: The post explains how scheduled, event-based, and self-reminding agent runs are managed in one place, along with the creator versus agent identity choice for unattended tasks.

  10. OpenCodeOfficialAI score60

    OpenCode Server has a code execution vulnerability in versions 1.14.30 through 1.18.21

    AIOpenCode warned that a code execution vulnerability affects OpenCode Server versions 1.14.30 through 1.18.21 and urged users to update to the latest version. The post credits @christophetd and the Datadog team for reporting the issue, with details linked in a Datadog Security Labs article.

    Why it matters: The post names the affected version range and the fix, which matters for anyone running OpenCode Server and deciding whether to upgrade now.

  11. Google Cloud · AI & Machine LearningOfficialAI score25

    Latin American midsize businesses adopt Google Cloud Gemini Enterprise to build AI agents

    AIAI adoption among Latin American small and medium-sized businesses has surged, with Google Cloud AI tool users growing 8x year-over-year across the region and 9x in Brazil. Companies such as AdGoat, Angelus, and BunkerDB are using Gemini Enterprise and Cloud infrastructure to automate content analysis, project management, and marketing workflows. BunkerDB reports cutting creative turnaround times from weeks to hours and reducing cost per lead by up to 25%.

  12. Liquid AI NewsletterOfficialAI score38

    Liquid AI optimizes its on-device context layer for Snapdragon processors and releases longevity models

    AILiquid AI says its Liquid Context on-device context layer is now optimized for Snapdragon processors using the Qualcomm Hexagon NPU, announced at Qualcomm's Snapdragon Summit. The company also released LFM2-1.2B-Longevity and LFM2-2.6B-Longevity, which it says match or outperform much larger frontier LLMs on longevity prediction, with LFM2-2.6B-Longevity more than 80% accurate on clinical age prediction. The open LongevityBench benchmark, with 17 tasks and 25,457 prompts, is available on Hugging Face.

  13. OpenBMBOfficialAI score34

    FIT-GGUF enables size-targeted mixed-precision quantization of MiniCPM5-2B

    AIDeveloper @Scorp1o_117 used FIT-GGUF to build four MiniCPM5-2B GGUF variants, ranging from about 1.14 GiB to 1.46 GiB, tuned to target file sizes or fidelity tiers. Instead of fixed presets, FIT-GGUF allocates precision tensor by tensor, with Quality, Balanced, Compact, and Mini options, and its generated files matched predicted sizes. Builds are evaluated with KL Divergence and Same-top metrics and are available on Hugging Face.

    Image from @OpenBMB's post
  14. Google · Innovation & AIOfficialAI score62

    Google's Project Suncatcher will test TPUs in orbit on a prototype satellite

    AIGoogle's Project Suncatcher will launch a prototype satellite on the Transporter-18 rideshare mission with SpaceX to test how its TPUs handle spaceflight. Initial ground tests showed the Trillium TPUs survived vibration and a radiation dose greater than a five-year space mission would deliver. Google says cooling with heat pipes and radiators and laser links between satellites in 2027 remain open engineering challenges.

    Why it matters: The source reports concrete radiation, vibration, and cooling test results for TPUs, showing what space-based AI compute still has to solve.

  15. TechNode · AINewsAI score34

    H3C Shifts AI Infrastructure Focus From More GPUs to Token Efficiency

    AIH3C argued at the 2026 Apsara Conference that AI infrastructure competition is shifting from adding GPUs to maximizing useful Tokens per GPU. The company showcased its UniPoD S80000 SuperPod, supporting 32 to 1,024 GPUs and scaling to 16,384, alongside switches for Scale-Up, Scale-Out, and Scale-Across interconnects. It also pitched its UniStor X20000 storage, which it says delivers up to 200GB/s bandwidth and cuts GPU waiting time by 30%.

  16. KrASIA · Big TechNewsAI score55

    Mind Lab launches Mint Recursive, a post-training platform for companies

    AIMind Lab unveiled Mint Recursive, a post-training and inference platform for industry use, alongside Macaron-V1.1, a model post-trained entirely on it. Macaron-V1.1 is a 752-billion-parameter model built from GLM-5.3 with four two-billion-parameter LoRA expert modules for chat, agents, coding, and generation. The platform is serverless and bills by token usage, and it collects feedback from models in use to support continued training.

  17. Lovable BlogOfficialAI score44

    Lovable Now Offers Free Chat for Planning and App Work

    AILovable now lets users chat for free to explore app ideas, review existing projects, and draft business materials before making changes. The chat can connect to tools like Notion, Granola, and Linear, and Free, Pro, and Business workspaces include a daily free chat allowance. Chats that generate images or video, or hand work off to Plan or Build, use credits as usual, and current chat pricing applies through October 31, 2026.

  18. Lovable BlogOfficialAI score80

    How Lovable's Chats connect conversations to agent work on projects

    AILovable describes how its Chats feature lets a workspace-level chat agent hand work to project builder agents and receive progress back. The design records each agent's history as an append-only, forkable trajectory, and passes messages through durable inboxes that activations wake. Agents can suspend at iteration boundaries and resume on freshly deployed nodes without killing long-running runs.

    Why it matters: The post details how trajectories, inboxes, and activations let agents share work and resume after deploys, useful for designing comparable agent systems.

  19. AI at MetaOfficialAI score42

    Meta unveils Muse Realtime Voice and Avatar with shared speech-token streaming

    AIMeta's Muse Realtime Voice generates speech tokens encoding both content and prosody, and Muse Realtime Avatar consumes that shared stream to produce streaming video. Using a fixed-length history as motion context keeps computation bounded regardless of conversation length while synchronizing voice, lip motion, and expressions.

    Image from @AIatMeta's post
  20. AI at MetaOfficialAI score22

    Meta distills 40-step video diffusion into a 2-step live streaming model

    AIMeta distilled a 40-step diffusion teacher using 3-way CFG, requiring 120 evaluations per video chunk, into an unguided 2-step causal student with a fixed-length KV cache. The student uses self-forcing to resist drift and keep near-teacher quality while needing 60x fewer evaluations, enabling instant responses in live video streaming.

    Image from @AIatMeta's post
  21. Jeff DeanXAI score20

    Jeff Dean highlights learning from computations in real rat neurons

    AIJeff Dean praises Alex Ksendzovsky and team at BioComputing Co for work that lets researchers learn from computations running in real rat neurons. The quoted post from Ksendzovsky says the company is partnering with Amazon to bring the technology to customers and is taking early-access requests.

  22. MiniMax (official)OfficialAI score34

    MiniMax-H3 video generation accelerated on AMD MI355X by Nunchux

    AINunchux runs MiniMax-H3 on AMD MI355X GPUs, generating 5 seconds of video in 1.3 seconds with up to 26.7x faster inference than SGLang on 8 GPUs. The stack supports streaming generation, letting users change prompts while the video plays. Free access to MiniMax-H3 through Nunchux is coming soon, with a waitlist open.

  23. Kling AI BlogOfficialAI score12

    Kling AI outlines six AI video limitations and workarounds for consistency and control

    AIKling AI's blog identifies six limitations of current AI video generation, including temporal consistency, character consistency across shots, unrealistic physics, long-form generation, fine details and text, and prompt control. It recommends workarounds such as reference images, shorter single-action clips, storyboards, and adding text or logos in post. The article says Kling VIDEO 3.0 and VIDEO 3.0 Omni offer reference-based subject consistency to help reduce these problems.

  24. LangChain BlogOfficialAI score50

    LangSmith Engine v2 adds red teaming and pre-validated agent fixes

    AILangChain released LangSmith Engine v2, an in-platform agent that scans production traces to detect agent issues and validates proposed fixes before human review. Engine v2 adds Red Teaming, currently in Private Beta for LangSmith Deployment users, which tests agents for weaknesses such as hallucinations and system-prompt violations before they reach production. Engine v2 is available in SaaS deployments for LangSmith Plus and Enterprise plans, with Self-Hosted support and BYOK for Engine coming later.

  25. LangChain BlogOfficialAI score50

    LangSmith Fine-Tuning and smithtune Turn Agent Trajectories Into Custom Models

    AILangChain launched LangSmith Fine-Tuning and smithtune, a CLI that turns LangSmith agent trajectories into fine-tuned models through dataset creation, training with Fireworks or Baseten, and evaluation in LangSmith. smithtune currently supports supervised fine-tuning, training models on recorded examples of good agent behavior by updating model weights. The tool lets teams train specialized models without building the data pipeline by hand.

  26. LangChain BlogOfficialAI score44

    LangSmith Launches Trajectories for Readable, Chronological Agent Session Views

    AILangChain has launched Trajectories in LangSmith, a chronological, conversational view that aggregates human, AI, and tool messages across an agent and its subagents. Trajectories work with traces from LangChain, LangGraph, Deep Agents, OpenAI and Claude agent SDKs, and coding agents like Codex, Claude Code, and Cursor. The feature is available now on all plans in the US.

Sep 23

Sep 23Wed
  1. Tencent HyOfficialAI score38

    Tencent Hunyuan studies batch-size scaling for LLM reinforcement learning efficiency

    AITencent Hunyuan extends classical critical-batch-size theory to online LLM reinforcement learning, where models generate their own training data. Across GRPO and PPO, learning-rate retuning preserves learning per response over a bounded range of batch sizes. On fixed hardware, larger batches raise PPO generation-stage throughput by up to 2.29×, and the best measured GRPO setup reaches the same validation target in 29% less time.

  2. OpenClaw🦞OfficialAI score10

    OpenClaw plugins can use optional models for structured decisions

    AISupporting plugins can call an optional model for structured choices, kept separate from chat. TypeSafe Jev sends supplied information to its hosted API and incurs normal charges, while ONNX offers local CPU options. Decision Models are off by default.

  3. OpenClaw🦞OfficialAI score12

    OpenClaw adds live meeting notes and saved transcript tabs

    AIOpenClaw lets users follow notes while a Google Meet, Teams, Zoom, or voice capture is still running. The saved transcript can be opened in its own tab, though transcription may lag. Generated notes use the user's configured model and incur its usual charges.

  4. OpenClaw🦞OfficialAI score19

    OpenClaw lets agents transfer files to a paired computer and back

    AIOpenClaw lets users send files to a paired computer where their agent works, then return finished files in chat. Memory and supported Skills can also be stored on that paired computer. Setup requires host configuration, permissions, and matching OpenClaw versions.

  5. WanOfficialAI score15

    Wan3.0 generates cinematic 1080p video with sound from one prompt

    AIAlibaba Cloud and Venice AI demonstrated Wan3.0 producing a finished scene from a single text prompt, with cinematic motion, 1080p output, and generated sound. The post presents the end-to-end workflow as a sign that the gap between an idea and a finished video is shrinking.

    Video from @Alibaba_Wan's post