Updated
#Deployment/Engineering
Updated
Aug 3
v0AI score12 v0AI score40 v0 launches new API for programmatic app building and deployment
AIv0 has introduced a new API giving programmatic access to its app-building capabilities. Developers can start chats from a prompt, repo, or ZIP, render a dev server preview, send follow-up messages, and deploy to Vercel.
Meituan LongCatAI score29 Meituan's LongCat launches international site with API and chat access
AIMeituan LongCat has officially launched its international site, offering overseas users a smoother experience. The API platform and chat interface are available at longcat.ai, with a Discord channel for support.
JetBrains AI BlogAI score52 JetBrains Built a Central CLI to Control Spiraling AI Tool Costs
AIJetBrains says its AI development expenses rose roughly 10x over six months as developers adopted three to five AI tools each. It built the JetBrains Central CLI, which routes third-party agent traffic through its AI platform so managers can set per-developer and team limits and view consumption reports. The CLI opened to early access on July 8 for anyone with JetBrains AI credits.
Manus BlogAI score38 Manus Adds ElevenLabs Connector for Chat-Based Audio Generation, Transcription, and Voice Apps
AIManus has launched an ElevenLabs connector that lets users generate speech, transcribe recordings, clone voices, and build audio apps through a single chat. Users connect their authorized ElevenLabs account via Integrations, and audio is processed within their own ElevenLabs environment according to its policies. Availability depends on users having an active ElevenLabs account, with capabilities tied to their ElevenLabs plan and credit balance.
Aug 2
OpenRouter BlogAI score40 OpenRouter Launches Ori Eval to Find the Best AI Model for Your App
AIOpenRouter has released Ori Eval, an agent-driven tool that runs your app's prompts against candidate models and returns a comparison table of catch rate, latency, cost per PR, and pass/fail results. The tool asserts on called tools and grades open-ended answers with an LLM judge, pinning the harness and model during each run. Its evals are code files that can run in CI to block regressions and re-run when new models ship.
Aug 1
Werner VogelsAI score22 Werner Vogels praises conversation with Clare Liguori on Kiro and agent support
AIWerner Vogels called his conversation with Clare Liguori an excellent discussion of developer support for agents and Kiro. The quoted InfoQ podcast covers moving agents from demo to production, including why extra if statements can hurt agent performance, achieving high accuracy and low cost with small models, and observability within agent hops.
Jul 31
DeepSeek · new models on Hugging FacePickAI score75 DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities
AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.
Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.
SkyworkAI score35 Skywork AI Hardware Family's first Skywork Note batch sells out in one week
AISkywork's first batch of its Skywork Note AI hardware device sold out one week after launch, prompting an accelerated rollout of the wider family, including the recording clip, the Recall pendant, and the TriRing AI ring. The company says the device is meant to capture real-world conversations and moments outside the screen, so users spend less time typing and more time away from it.
DeepSeekAI score38 DeepSeek-V4-Flash-0731 API upgrade keeps preview architecture and size
AIDeepSeek says DeepSeek-V4-Flash-0731 keeps the same model architecture and size as the preview version. Today's upgrade applies only to the DeepSeek-V4-Flash API, while the DeepSeek-V4-Pro API and App/Web models remain unchanged for now. The official DeepSeek-V4-Pro release is coming soon.
DeepSeek API NewsPickAI score67 DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks
AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.
Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.
Jul 30
Jeff DeanAI score38 Jeff Dean thanks Diana Hu after Startup School conversation at Chase Center
AIJeff Dean, Google's Chief Scientist, thanked YC partner Diana Hu for an engaging conversation at Chase Center last weekend, his first in a basketball arena. The post is a brief acknowledgment, with the surrounding context describing a Startup School 2026 discussion on AI inference hardware, the origins of TPUs, and advice for founders.
Mark ChenAI score62 OpenAI cuts GPT-5.6 Luna and Terra API prices and adds Fast mode to Sol
AIOpenAI cut API prices for two GPT-5.6 models, with Luna down 80% to $0.20 per million input tokens and $1.20 per million output tokens. Terra drops 20% to $2 input and $12 output per million tokens. GPT-5.6 Sol gains a Fast mode in the API that offers up to 2.5x the speed for 2x the price at the same intelligence level.
Thinking MachinesAI score38 Inkling and Inkling-Small now available on Tinker with discount
AIThinking Machines announced that its Inkling and Inkling-Small models are both available on Tinker with a limited-time discount. All Tinker models can now also be chatted with on the Tinker Playground.
IdeogramAI score38 P-Image-Ideogram launches via API and partner platforms broadly
AIIdeogram has made P-Image-Ideogram available now through its API and a long list of partner platforms, including ComfyUI, Runware, Replicate, Leonardo.Ai, Picsart, Cloudflare, Magnific, Gamma, and Together AI. The company says the rollout aims to broaden access to frontier-quality image generation.
IdeogramAI score20 Ideogram's P-Image-Ideogram Showcases Photorealism and Text Rendering
AIIdeogram highlighted four examples of its P-Image-Ideogram model demonstrating photorealism, text rendering, and style versatility. The post claims the model offers the best quality for the price and supports JSON prompting and layout control.
IdeogramAI score23 Ideogram unveils P-Image-Ideogram for production pipelines with quality modes
AIIdeogram's P-Image-Ideogram offers four quality modes (Very low, Low, Medium, High), native 1K and 2K generation, and a range of aspect ratios. Latency runs from 3s to 8s, letting users choose their point on the speed-quality curve.
IdeogramAI score43 Ideogram and Pruna launch P-Image-Ideogram image model family
AIIdeogram introduced P-Image-Ideogram, a family of image models co-developed with Pruna, offering a quality-speed-cost trade-off across four quality modes. The models generate native 1K and 2K images starting at $0.003 per image and are available now on the Ideogram API and partner platforms.
Jul 29
Fireworks AI BlogAI score54 Fireworks tests whether LoRA or full fine-tuning gaps come from data, learning rate, or rank
AIFireworks AI ran controlled SFT experiments on Qwen3.5-9B comparing LoRA with full parameter fine-tuning across three synthetic verifiable tasks. The post argues that a FullFT advantage can come from data coverage, learning-rate tuning, or adapter rank, and it recommends testing these in that order before switching methods. Under a fixed multi-task budget, FullFT kept a 4.29-point lead over the best LoRA recipe tested, while matched data exposure favored LoRA.
MicrosoftAI score34 Microsoft reports FY26 Q4 revenue of $90 billion, Copilot tops 30 million seats
AIMicrosoft reported FY26 Q4 earnings with $90 billion in revenue, and Azure and other cloud services revenue grew 43% on strong demand across the platform. Microsoft 365 Copilot reached more than 30 million paid seats, and Satya Nadella said Azure revenue surpassed $100 billion for the first time this year.
Ahmad Al-DahleAI score52 Ahmad Al-Dahle argues AI capex is both short on compute and overbuilt
AIAhmad Al-Dahle argues that AI infrastructure faces both a compute shortage and overbuilding, with the four largest hyperscalers planning roughly $725 billion of capex in 2026, up 77 percent from last year. He describes a "mutually assured construction" dynamic in which every well-capitalized player buys the same insurance against falling behind, so the industry overbuilds by construction.
Hugging FaceAI score22 Hugging Face adds Sign-In OAuth for apps with repos, Buckets, and Jobs
AIHugging Face is encouraging developers to add a "Sign-In with Hugging Face" OAuth button to their websites and apps. Users could then share their email, create model and dataset repos, store data in Buckets, or start GPU-backed Jobs.
Berkeley AI ResearchAI score44 K-Search Adapts CUDA Kernel Expertise to Apple Silicon MLX Backend
AIBerkeley AI Research extended the K-Search evolutionary kernel framework with an MLX backend and a CUDA-to-MLX translation layer, letting it adapt existing CUDA kernels for Apple Silicon. The team reports a 0.97x speedup relative to the native MLX Attention kernel and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel. The method uses Gemini 3.5 Pro Preview to both reason about optimizations and write candidate kernels.
Jul 28
Augment Code BlogAI score39 GPT-5.6 Sol Becomes Augment Cosmos's Default Model for Token Efficiency
AIAugment Code has made GPT-5.6 Sol the default model in Cosmos, choosing it as the most token-efficient model to clear its pass-rate floor for long-horizon software engineering tasks. The company ranks models by cost per task rather than list price per million tokens, since retries on failed steps add token spend. Users can still select any model, and the default will change as more token-efficient models emerge.
Fireworks AI BlogAI score46 Fireworks AI Shows Low-Cost Fine-Tuning Lifts Domain Embedding Retrieval
AIFireworks AI describes fine-tuning Qwen3-Embedding-8B on private (query, positive) pairs using bidirectional InfoNCE loss through its Training SDK, then serving the model via an OpenAI-compatible embeddings endpoint. The post reports that around 150 training steps was enough, that rank-32 LoRA landed within about one point of full-parameter fine-tuning, and that gains were largest where the base model struggled, while tasks like CoSQA and FiQA2018 showed flat results.
Cognition Blog (Devin, Windsurf)AI score28 LTM Partners with Cognition to Deploy Devin for Cybersecurity Risk Reduction
AILTM has partnered with Cognition to deploy Devin, the AI software engineer, through BlueVerse RightLogic, a managed, outcome-based service that clears customers' vulnerability backlogs. RightLogic is designed to clear 80 percent of an enterprise's CVE backlog, up from the 60 percent previously delivered, and will focus first on banking, financial services, and insurance. The service is the first of five joint offerings the companies plan to bring to market.
JetBrains AI BlogPickAI score60 Ponytail Skill Cuts Claude Code Costs 10% But Not the Advertised 54%
AIJetBrains tested the ponytail skill for Claude Code across 80 paired tasks and found a median 10.3% cost reduction, with p=0.004. Code written fell about 15% median versus the advertised 54%, reaching 31% on larger builds and little on already-lean tasks. No quality difference was detected, and the skill only self-activated when its ruleset was injected by a plugin hook.
Why it matters: The benchmark separates advertised savings from measured results and shows the code cut depends on how much the baseline agent over-builds.
Aman SangerAI score39 Cursor launches ₹649/month Start plan in India with Grok 4.5 and Composer
AICursor says its India user base tripled in the past year and that Indian users make more agent requests per user than any other country. It is launching Cursor Start, a ₹649/month plan with generous access to Grok 4.5 and Composer for Indian developers.
Jul 27
Sequoia CapitalAI score24 Cyera to Acquire Oasis Security to Combine Data and Identity Security for AI
AICyera is joining with Oasis Security, which builds agentic access management for non-human identities such as API keys, service accounts, OAuth tokens, and agent credentials. The combination pairs Cyera's knowledge of where sensitive data lives with Oasis's visibility into which identities and agents can reach it. Sequoia Capital, which backed both companies since their Series A rounds, says the pairing covers the full path an AI agent takes through an enterprise.
Liquid AI BlogAI score49 Liquid AI Releases LFM2.5-Encoders for Fast Long-Context Encoding on CPU
AILiquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, bidirectional encoders built on the LFM2 hybrid architecture and available on Hugging Face. They support an 8,192-token context and are designed for fine-tuning on classification and token-level tasks. On CPU, LFM2.5-Encoder-230M is the fastest model tested from 1K tokens up, running about 3.7x faster than ModernBERT-base at 8,192 tokens.
KimiAI score52 Kimi K3 launches on Together AI as a Day 0 partner
AIKimi K3 is now available on Together AI, which is a Day 0 launch partner for the model. Together AI offers developers immediate access to K3 through high-throughput inference aimed at coding agents and production workloads.
MetaAI score26 Meta's SAM3 and DINO help Pittsburgh researchers build advanced wheelchair
AIUniversity of Pittsburgh researchers used Meta's SAM3 and DINO models to help build one of the most advanced wheelchairs in the world. The post provides no further technical details on performance or availability.
KimiAI score38 Kimi K3 now available on DigitalOcean Serverless Inference
AIMoonshot AI's Kimi K3 is now available on DigitalOcean's Serverless Inference, letting developers start building in minutes. DigitalOcean describes K3 as supporting a 1M-token context, native vision, and multi-hour agentic tasks, and it is accessible through the Inference Router.
KimiPickAI score65 Kimi K3 becomes available on Nebius Token Factory via API
AIKimi K3 is now available on Nebius Token Factory, which is named a Day 0 launch partner, through an OpenAI-compatible API and console. The quoted post says Artificial Analysis scores the open-weight model at 57 on its Intelligence Index, two points behind GPT-5.6 Sol (max), and lists up to 1M tokens of context.
Why it matters: The source names the cloud access route and an Artificial Analysis score of 57, letting readers compare Kimi K3 against GPT-5.6 Sol.
KimiAI score62 Kimi K3 launches on Fireworks with day-0 inference and fine-tuning
AIKimi K3 is available on Fireworks from day 0 for inference and training, hosted in the US with zero data retention. The author says users can deploy and fine-tune the 2.8T-parameter model with a few clicks.
KimiAI score35 Kimi K3 launches day 0 on Baseten's Model APIs
AIMoonshot AI's Kimi K3 is available from day zero through Baseten's Model APIs, offering fast and reliable access. Baseten is named as the launch partner for bringing K3 to more users.
KimiAI score47 Kimi K3 launches with Modal as Day 0 partner for faster inference
AIKimi K3 is available on Modal as a Day 0 launch partner, with Modal training a custom DFlash speculator for the model's architecture. The speculator delivers faster inference with no quality loss, according to Kimi. Modal describes K3 as a 3T-class open model that is the most capable open model it has worked with.
KimiAI score31 Moonshot AI open-sources MoonEP, a communication library for MoE workloads
AIMoonshot AI has open-sourced MoonEP, a high-performance communication library for distributed MoE workloads. The library is designed to reduce communication overhead in expert-parallel training and inference at scale. The code is available on GitHub under MoonshotAI/MoonEP.
KimiAI score38 Kimi and kvcache-ai open-source AgentENV for scalable agent environments
AIMoonshot AI's Kimi, in collaboration with kvcache-ai, has open-sourced AgentENV, a distributed system for running agent environments at scale. Its components power agentic RL training for Kimi K3, supporting fast snapshot, resume, and fork for large-scale parallel agent workflows. The project is available on GitHub at
KimiAI score40 Moonshot AI open-sources FlashKDA, a CUTLASS-based Kimi Delta Attention kernel
AIMoonshot AI has open-sourced FlashKDA, a high-performance CUTLASS-based implementation of Kimi Delta Attention kernels. It delivers a 1.72×–2.22× prefill speedup over the flash-linear-attention baseline on H20 and works as a drop-in backend for flash-linear-attention.