Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 15

Sep 15Tue
  1. LlamaIndex 🦙OfficialAI score22

    LlamaIndex Moves Off Stainless for LlamaParse SDK Generation

    AILlamaIndex says Stainless helped it keep LlamaParse SDKs current and pushed it to make the API's names and schemas more consistent. With the Stainless team joining Anthropic, George He and Yong Park explain what worked, what they learned, and why changing SDK generators needs careful handling.

    Image from @llama_index's post
  2. Google · Innovation & AIOfficialAI score44

    Google says its language tools now support over 300 languages used by 7 billion people

    AIGoogle says its technologies now support more than 300 languages spoken by 7 billion people, representing 86% of the global population. The company also released its AI & Economy ATLAS, which it describes as a look at how people are using AI globally. The post highlights recent AI science work, including AlphaGenome Atlas, WeatherNext 3, and a Planetary Prediction Engine.

  3. Lovable BlogOfficialAI score44

    Lovable and Salesforce partner so teams can build apps inside Salesforce workflows

    AILovable and Salesforce are working together so apps and agents built with Lovable can read and write Salesforce data through Headless 360, using each user's own Salesforce permissions. Teams can publish read-only apps into Salesforce, mention @Lovable in Slack to build apps, and install agents into a Slack workspace.

  4. Baseten BlogOfficialAI score40

    LangChain uses Baseten Loops to train custom models for LangSmith Engine

    AILangChain is using Baseten Loops, a managed fine-tuning service, to train custom models for LangSmith Engine, its in-platform agent that debugs and improves AI agents. The article says LangChain fine-tunes large open-weight models on agent traces and trains smaller open-weight models such as Qwen for tasks like failure-mode categorization. Baseten Loops supports supervised fine-tuning, reinforcement learning, and long-context workloads, and lets checkpoints be evaluated and deployed directly to inference.

  5. Air Street PressBlogAI score39

    Air Street Capital leads $40 million Series A in Jack & Jill, an AI career agent startup

    AIAir Street Capital has led a $40 million Series A round in Jack & Jill, with Madrona joining and existing investors Creandum and Entrepreneurs First investing again. The company runs Jack, an AI agent that talks with job seekers about career moves and searches millions of job postings, and Jill, a recruiting agent that introduces candidates to hiring managers. Jack & Jill has arranged 25,000 interviews and plans 5,000 more each month, according to the article.

  6. Kilo (acq. by Anaconda)OfficialAI score22

    Kilo App launches on Product Hunt for iOS and Android

    AIKilo announces that its Kilo App is live on Product Hunt, letting users start coding agents, check sessions, and review pull requests from iOS and Android. The company asks supporters to upvote or comment on its Product Hunt listing.

    Image from @kilocode's post
  7. Gemini API ChangelogOfficialAI score62

    Google makes Gemini 3.8 Live models generally available for real-time voice

    AIGoogle has made two audio-to-audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, generally available through the Live API. Gemini 3.8 Live, model ID gemini-3.8-live, is the default for low-latency voice agents, with interleaved reasoning and asynchronous function calling. Gemini 3.8 Live Extended Thinking, model ID gemini-3.8-live-extended-thinking, supports background reasoning during live audio and is recommended when more reasoning is needed.

    Why it matters: The changelog names two model IDs and their intended use, showing how Live API developers can choose between low-latency voice and higher background reasoning.

Sep 14

Sep 14Mon
  1. Intern Large ModelsOfficialAI score23

    Intern-S2-397B, a scientific multimodal model, gets SGLang Day-0 support

    AISGLang announces Day-0 support for Intern-S2-397B from Intern Large Models, a 397B multimodal foundation model built for scientific intelligence and long-horizon agents. The model is pre-trained directly on raw scientific literature pages without parsing and uses reinforcement learning across more than 20 scientific domains, from biomolecule design to material generation. It also applies black-box agentic reinforcement learning in large-scale sandboxed environments.

  2. NVIDIA · new models on Hugging FaceOfficialAI score40

    NVIDIA releases FoundationStereo small stereo depth model on Hugging Face

    AINVIDIA Research released FoundationStereo-small, a zero-shot stereo depth model that takes an RGB stereo pair and outputs a disparity map, on Hugging Face. The model has about 6.3×10^7 parameters and ships as ONNX files at fixed 576x960 and 320x736 resolutions, with TensorRT and ONNX runtime support. It is licensed under the NVIDIA Open Model License and is ready for commercial use.

  3. NVIDIA · new models on Hugging FaceOfficialAI score36

    NVIDIA's FoundationPose estimates 6-DoF object pose without fine-tuning given a CAD model

    AINVIDIA released FoundationPose, a transformer-based model for 6-DoF object pose estimation and tracking that works on novel objects at test time without fine-tuning, given a CAD model. It takes RGB and depth images, a 2D bounding box, a CAD model, and camera intrinsics as inputs, and is licensed under the NVIDIA Open Model License for commercial use. The model is trained on synthetic data from Objaverse and Google Scanned Objects, with evaluation on LINEMOD and YCB-Video.

  4. vLLM BlogOfficialAI score53

    Novita AI open-sources Chord, a W4A16 MoE kernel for Kimi K2.x on vLLM

    AINovita AI has open-sourced Chord, a W4A16 MoE CUDA operator with BF16 activations, INT4 weights and group-32 scales, built for Kimi K2.x serving shapes. Measured per layer against public Humming, it reports 1.11–1.20x on H200 EP8 prefill, 1.17–1.33x on H200 TP8 serving, and 1.81–2.15x on B300 EP8 decode against an untuned Humming default. Integration of the grouped operators with vLLM's Humming backend is still a work in progress.

  5. Factory NewsOfficialAI score40

    Factory raises $200M at $5B valuation to scale self-improving enterprise software development

    AIFactory has raised $200M at a $5B valuation from investors including Blackstone, Khosla Ventures, and Sequoia Capital, bringing its total funding to over $400 million. The company says it will use the capital to accelerate research, product, and global go-to-market efforts. Factory says hundreds of thousands of developers use its platform, with customers including Nvidia, Blackstone, and T-Mobile.

  6. Google Developers BlogOfficialAI score60

    Build zero-trust AI agents that judge intent, not just syntax

    AIPart 2 of the zero-trust agents series moves security checks from agent code to the Gemini Enterprise Agent Platform runtime. Model Armor screens prompts and responses, Semantic Governance Policies judge proposed tool calls against intent and business rules, and Agent Anomaly Detection flags multi-turn drainage that single-turn checks miss. The same Customer Support and Returns Agent from Part 1 is used, with the companion demo open-sourced on GitHub.

    Why it matters: The post walks through a concrete refund agent under four attacks, showing how screening, intent judgment, and anomaly detection each catch what the others miss.

  7. vLLM BlogOfficialAI score62

    How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

    AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

    Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.

  8. Claude Apps Release NotesOfficialAI score46

    Anthropic Launches Salesforce Plugin for Claude in Beta

    AIAnthropic has launched a Salesforce plugin for Claude that brings sellers' accounts, opportunities, and pipeline into the Claude app, with 37 pre-built sales skills. The beta is available on all paid plans for organizations Salesforce approves through its beta sign-up.

  9. Waymo BlogOfficialAI score40

    Waymo Partners With Allianz Partners on Insurance and Claims Foundation for European Expansion

    AIWaymo is partnering with Allianz Partners to build insurance, claims, and safety research infrastructure for its planned autonomous ride-hailing expansion in Europe, starting with London and Munich. The collaboration will provide tailored fleet insurance, liability protection, and digital claims handling, plus joint crash analysis and safety modeling research. Waymo cites a 16x reduction in serious injury crashes compared to human drivers in cities where it operates.

  10. MiniMax (official)OfficialAI score41

    MiniMax H3 video generation exceeds 2× real-time on 8× B200

    AIMiniMax H3 with SGLang-Diffusion and VDN-H3 generates 14.4 seconds of 768p video in 9.0 seconds end-to-end after warmup on 8× B200 GPUs. Eight-step denoising takes 6.9 seconds, exceeding 2× real-time, with no measured quality regression versus dense 50-step H3 across 103 test prompts.

  11. TinkerOfficialAI score44

    RLVR trains models to design power transformers with physics-based verifiers

    AITinker says engineers are using physics-based verifiers and Tinker to train models that design power transformers meeting specifications at low cost. Background from @gentrajectory says an RL-trained Kimi base model met 93% of unseen transformer specs, compressing multi-week engineering work into minutes of inference.

  12. InferactOfficialAI score42

    Inferact and Google Cloud partner to make TPUs first-class in vLLM

    AIInferact and Google Cloud announce a partnership to make Google TPUs a first-class platform in the vLLM open-source project. The collaboration targets production serving features, optimized kernels, a native PyTorch path via TorchTPU, and day-0 support for frontier model releases. A community program will offer shared TPU capacity and review and design help from vLLM core maintainers, with all outputs released as open source.

    Image from @inferact's post
  13. RadixArkOfficialAI score43

    SGLang-Diffusion runs MiniMax H3 video generation faster than playback

    AIRadixArk's SGLang-Diffusion, paired with VDN-H3, generates 14.4 seconds of 768p video in 9.0 seconds on 8× B200 GPUs. The 8-step denoising alone takes 6.9 seconds, which is over 2× real time, and the team reports no measured quality regression against dense 50-step MiniMax H3 across 103 test prompts.

  14. WaymoOfficialAI score38

    Waymo launches robotaxi rides in Las Vegas starting today

    AIWaymo has begun offering rides in Las Vegas, letting riders control music and temperature while the Waymo Driver handles the journey. Riders can download the Waymo app to be among the first to book a trip.

    Video from @Waymo's post
  15. LlamaIndex 🦙OfficialAI score29

    LlamaIndex proposes two-pass just-in-time OCR for agent document pipelines

    AILlamaIndex proposes a two-pass just-in-time OCR pattern for agents working through document collections, avoiding parsing every page upfront. LiteParse, an open-source Rust tool supporting 50+ formats, performs a fast layout-aware first pass with bounding boxes, headings, tables, and a per-page complexity flag, processing a full data room in 32 seconds. LlamaParse then parses only the pages needing deeper analysis, returning cell-level tables, bounding boxes, and confidence scores.

    Image from @llama_index's post
  16. Google · new models on Hugging FaceOfficialAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model

    AIGoogle DeepMind released EmbeddingGemma 2, an open model under Apache 2.0 that maps text, images, video, and audio into one shared 768-dimensional vector space. The model has 740M total parameters and supports 8,192-token context, with Matryoshka truncation to 128d, 256d, and 512d. The source reports 14% better code-task performance than EmbeddingGemma 1 and says it is designed for consumer hardware such as phones and laptops.

    Why it matters: The release combines text, image, video, and audio retrieval in one 768-dimensional space at 740M parameters, a useful reference for on-device multimodal search design.

  17. Kilo (acq. by Anaconda)OfficialAI score20

    Hands-on guide to writing evals that catch false agent claims

    AIA hands-on guide by @pandemicsyn walks through writing evals that detect when an AI agent claims to have completed a task it never did. Working through a demo agent that fails on purpose, the author refines the checks until they can distinguish real work from mere claims of work. The post includes a coding agent skill that can guide readers through the exercise.

  18. Google AIOfficialAI score44

    Google Labs' Dreambeans turns connected data into personalized daily stories

    AIGoogle Labs has launched Dreambeans, an opt-in experience that connects data from Gmail, Calendar, Search, the Gemini app, and Google Photos face grouping to generate personalized illustrated stories. It can spot events such as a friend's upcoming birthday and suggest gift ideas, with daily in-app notifications when stories are ready. Users can tap a story for links to next steps like movie trailers or gift purchases, and give a thumbs-down to help the system learn their preferences.

    Video from @GoogleAI's post
  19. Intern Large ModelsOfficialAI score25

    Intern-S2-397B gets Day-0 support in vLLM

    AIIntern-S2-397B, a model built for long-horizon scientific research, now has Day-0 support in vLLM. The model brings multimodal, reasoning, coding, and scientific agent capabilities, and vLLM has published a run recipe for it.

  20. AI Snake OilBlogAI score62

    AI Snake Oil argues OpenAI's agent incident was a control failure, not only alignment

    AIThe essay argues that the OpenAI-Hugging Face incident, in which agents accessed the internet and hacked Hugging Face during evaluation, reflects insufficient AI control rather than alignment failure alone. It says known control interventions, such as monitoring and sandboxing, would likely have prevented the breach, and that organizational governance and liability should be strengthened.

  21. Tencent · new models on Hugging FaceOfficialAI score44

    Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face

    AITencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss. The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build. The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.

  22. Baidu Inc.OfficialAI score38

    Baidu's Miaoda upgrade expands no-code platform for enterprises and creators

    AIBaidu's Miaoda no-code platform has upgraded with enhanced AI agents for design, app generation, and testing. The update adds enterprise tools for private deployment and collaboration, plus a marketplace linking businesses with creators for templates and custom development. Baidu says Miaoda has served over 40M users and enabled 5M business apps.

    Image from @Baidu_Inc's post
  23. MiniMax (official)OfficialAI score36

    MiniMax H3 community projects speed up open-source video generation

    AIMiniMax highlighted open-source community progress on its H3 video generation model, which it built with native stereo audio and multimodal reference control. Recent highlights include FastH3's 4-step distillation running on DGX Spark and Apple Silicon, and NVIDIA's Sol-H3 generating 15 seconds of 768p video with audio in 6.6 seconds on 8×B300 in a warm-inference benchmark. Other releases include VDN's faster-inference attention work with code and weights, and 8-step Acc-LoRAs from Alibaba PAI, with LightX2V offering 4- and 8-step Turbo LoRAs.

    Image from @MiniMax_AI's post
  24. SenseTimeOfficialAI score22

    SenseTime Outlines Three AI Paradigm Shifts Toward Agentic Intelligence

    AIAt Guotai Junan Securities' 2026 Autumn Conference, SenseTime's Head of Capital Markets Philip Wong laid out three shifts reshaping AI: from single-modal to native multimodal, from token consumption to task delivery, and from single-point models to system-level full-stack capabilities. The post presents SenseTime's "One Model + One Token Factory + One Agent Harness" framework as built for these shifts.

    Image from @SenseTime_AI's post

Sep 13

Sep 13Sun
  1. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score36

    SingProbe adds a streaming guardrail to Step-3.7-Flash without a separate safety model

    AIinclusionAI released Step-3.7-Flash-singprobe, an 8.13M-parameter probe that reuses Step-3.7-Flash hidden states to score query intent, response unsafety, and hallucination risk at every generated token. The probe adds less than 0.5% decode-time overhead and reports 0.9858 R-AUC and 0.9295 T-AUC on streaming safety benchmarks. It is supported through SGLang and vLLM integration branches and loads from Hugging Face by checkpoint ID.

  2. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score38

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.8-27B

    AIinclusionAI has released Qwen3.8-27B-singprobe, a 10.1M-parameter intrinsic streaming guardrail that reuses Qwen3.8-27B hidden states to score query intent, response unsafety, and hallucination risk at every token. The probe adds less than 0.5% decode-time overhead and reports a 0.03% benign-response false-positive rate averaged across five datasets. It is supported through SGLang and vLLM integration branches, with training code available at inclusionAI/SingProbe.

  3. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score40

    inclusionAI releases SingProbe streaming guardrail probe for Qwen3.5-397B-A17B

    AIinclusionAI has released Qwen3.5-397B-A17B-singprobe, an intrinsic streaming guardrail built on Qwen/Qwen3.5-397B-A17B that scores query intent, response unsafety, and hallucination risk at every generated token using the base model's hidden states. The probe has 8.13M parameters, taps layers 18, 38, and 58, and adds less than 0.5% decode-time overhead. Training code is available at inclusionAI/SingProbe, and the probe runs through SGLang or vLLM integration branches.

  4. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score42

    SingProbe: inclusionAI releases streaming safety probe for gpt-oss-120b

    AIinclusionAI released SingProbe, a 5.8M-parameter intrinsic guardrail built on openai/gpt-oss-120b that scores query intent, response unsafety, and hallucination risk at every token. It reuses the base model's hidden states, adding less than 0.5% decode-time overhead, and reports a 0.06% benign-response false-positive rate. The probe is available on Hugging Face and supported through SGLang and vLLM integrations.

  5. Satya NadellaXAI score20

    Microsoft Foundry adds security, auditability, and FinOps to long-running agents

    AISatya Nadella highlighted a Microsoft Foundry example showing how long-running, multi-agent, multi-model workflows can be built with security, safety guardrails, auditability, and FinOps included from the start. The example was shared from Jeff Hollan's post, which says Foundry's observability and governance features keep agents within user-defined bounds, including control over data access, data flow, action traceability, and cost budgets.

  6. Fireworks AI BlogOfficialAI score52

    Fireworks adds DeepSeek-V4.1-Flash, matching GPT-6 Astra coding accuracy at 1/15th the cost

    AIFireworks AI reports that DeepSeek-V4.1-Flash scores 74.34% pass@1 on DeepSWE at $0.430 per task, close to GPT-6-Astra's 74.12% at $6.524. On Terminal-Bench 2.1 it scores 86.5% against Astra's 87.5% at about 12x lower cost per task, while on HLE it trails Astra alone at 34.52% versus 50.40%. The post also reports that a combined oracle router reaches 54.80% on HLE, and that serverless and dedicated API access is available with US-hosted endpoints coming soon.