Skip to contentSkip to stories

Updated

#Image generation

Showing low-relevance items too. Hide low-relevance items

Sep 20

Sep 20Sun
  1. WanOfficialAI score22

    Wan 3.0 turns a single photo into an animated scene swap

    AIWan 3.0 lets users place their own photo into any scene and animate it using the Peel-Off Sticker skill on wan.video. The post presents this as a single-photo workflow but gives no further technical details.

    Video from @Alibaba_Wan's post

Sep 18

Sep 18Fri
  1. Google AIOfficialAI score47

    Google's weekly recap: Gemini 3.8 Live, Dreambeans, CC, and more

    AIGoogle's weekly recap covers Gemini 3.8 Live and 3.8 Live Extended Thinking, described as its most advanced live dialogue audio models yet. It also notes Dreambeans, a GoogleLabs experiment curating daily personalized stories, is now generally available, and that CC has expanded into a shared agent for household coordination. Google Pics, a Workspace tool for generating and co-creating images, is now GA, alongside AlphaGenome Atlas, DeepMind's interactive genomics discovery platform.

  2. Mustafa SuleymanXAI score22

    Mustafa Suleyman claims best image generation quality-price performance

    AIMustafa Suleyman, owner of the source account associated with Microsoft and Copilot, says the post claims the best image generation quality-price performance in the world. The post itself gives no specific model, price, or benchmark figures. Background from Artificial Analysis says Muse Image, MAI-Image-2.6, and GPT Images 2.5 recently shifted text-to-image price and speed frontiers.

Sep 17

Sep 17Thu
  1. World LabsOfficialAI score32

    World Labs' Atlas generates real-time flythrough from 32 input images

    AIWorld Labs says its Atlas model turns 32 input images into a real-time flight through NVIDIA's Voyager headquarters. Trained on NVIDIA Blackwell GPUs, Atlas uses the images as 3D spatial context to generate new views with pixel-perfect camera control.

    Video from @theworldlabs's post
  2. Daniel HanXAI score44

    Unsloth Desktop adds multi-user accounts and faster GRPO training

    AIUnsloth Desktop now supports multi-user accounts, alongside a revamped Docker image and custom Jupyter Notebook with custom themes, titles, and expandable cells. The update adds RDNA1+2 support, ARM64 Windows CUDA support, faster GRPO, and FP8/INT8 image diffusion support for 2x faster inference.

  3. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score46

    Ming-Image-0.1-Design-Layer splits flattened design images into RGBA layers

    AIinclusionAI has released Ming-Image-0.1-Design-Layer on Hugging Face, a model that decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan. The model runs at 1024 resolution (512 for faster processing) with 12 sampling steps, a CFG scale of 2.0, and BF16 precision on one CUDA GPU with 80 GiB VRAM. It is released under the MIT License.

  4. inclusionAI (Ant Ling) · new models on Hugging FaceOfficialAI score42

    inclusionAI releases Ming-Image-0.1-Design, a 6B text-to-image model for text-rich designs

    AIinclusionAI has released Ming-Image-0.1-Design, a 6B text-to-image model for UI, infographics, and posters that outputs RGBA images with transparent backgrounds. The model is available on Hugging Face and ModelScope under the MIT License. It runs at 2048 x 2048 with 12 sampling steps and a CFG scale of 1.0, validated on one CUDA GPU with 80 GiB VRAM.

Sep 16

Sep 16Wed
  1. Kling AIOfficialAI score3

    Kling AI shares a link to a full film

    AIKling AI's post links to a full film hosted at watch.fountain0.com. The post itself gives no further details about the film's content, production, or release.

  2. BAAI · new models on Hugging FaceOfficialAI score34

    BAAI and Peking University release Brainμ-Spike spike camera image reconstruction model

    AIPeking University's Yu Zhaofei team and the Beijing Academy of Artificial Intelligence (BAAI) released Brainμ-Spike, a small convolutional network for spike camera image reconstruction that is paired with the Brainμ model. The package includes weights, inference scripts, and evaluation tools, but the base large model and LoRA weights are not yet released, so the full generation pipeline cannot run from this repository alone.

Sep 15

Sep 15Tue
  1. Stability AIOfficialAI score25

    Stability AI's team presents color consistency research at ECCV

    AIStability AI's interactive research team presented new work at the 19th European Conference on Computer Vision (ECCV) aimed at keeping colors consistent across shots and AI-generated reference photos. The post frames color consistency as a persistent friction point in production.

    Image from @StabilityAI's post
  2. MiniMax Design (H3)OfficialAI score16

    Hailuo AI's MiniMax H3 brings GPT-6 Astra-generated set to life

    AIHailuo AI, owned by MiniMax, promotes its MiniMax H3 model as a way to create an army of characters without casting or catering costs. A creator, @koldo2k, says they used GPT-6 Astra-generated work to build a set and story, then animated it with MiniMax H3.

Sep 14

Sep 14Mon
  1. MiniMax Design (H3)OfficialAI score9

    MiniMax Hailuo AI Bestiary contest closes tonight with $8,000 prize

    AIThe Hailuo AI Bestiary contest closes tonight, September 14 at 23:59 PT, with $8,000 cash, 200,000 Credits, and ten Audience Choice awards. Entrants must submit a 30-second-or-longer short film generated with MiniMax H3 on any myth, era, or world.

  2. MiniMax (official)OfficialAI score36

    MiniMax H3 community projects speed up open-source video generation

    AIMiniMax highlighted open-source community progress on its H3 video generation model, which it built with native stereo audio and multimodal reference control. Recent highlights include FastH3's 4-step distillation running on DGX Spark and Apple Silicon, and NVIDIA's Sol-H3 generating 15 seconds of 768p video with audio in 6.6 seconds on 8×B300 in a warm-inference benchmark. Other releases include VDN's faster-inference attention work with code and weights, and 8-step Acc-LoRAs from Alibaba PAI, with LightX2V offering 4- and 8-step Turbo LoRAs.

    Image from @MiniMax_AI's post
  3. Sherwin WuXAI score40

    GPT Image 2.5 now live in ChatGPT with better editing consistency

    AIGPT Image 2.5 is now available in ChatGPT, and the post says it keeps consistency well while editing images. It also claims state-of-the-art results on all image leaderboards, and suggests users who saw faces shift in GPT Image 2 edits try again.

    Video from @sherwinwu's post

Sep 13

Sep 13Sun
  1. Qwen · new models on Hugging FaceOfficialAI score67

    Qwen releases open-source Qwen-Image-2.1 for generation and editing

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with 7B parameters in its visual generation component. The model can generate regular or transparent RGBA images, supports up to 10 reference images for editing, and is licensed under the Qwen Research License Agreement.

    Why it matters: The source specifies the 7B visual component, transparent RGBA output, and up to 10 reference images, which helps readers judge its fit for generation and editing workflows.

Sep 8

Sep 8Tue
  1. Ian Johnson 🔬🤖XAI score23

    Ian Johnson builds a font generator from letter-cluster embeddings

    AIIan Johnson (@enjalot) built a font generator after finding a cluster for each letter of the alphabet in his dataset, with Astra helping write the code. The tool is available as a Hugging Face Space and on GitHub, and the dataset includes SigLIP2 embeddings that allow concept search and clicking a result to jump to similar blocks.

    Video from @enjalot's post
  2. ReplicateOfficialAI score46

    GPT Image 2.5 Launches on Replicate for Precise Image Editing

    AIOpenAI's GPT Image 2.5 is now available on Replicate, offering precise image editing, sharper details, and higher consistency across multi-turn edits. Two model pages are provided, openai/gpt-image-2.5-sunburst and openai/gpt-image-2.5-flare.

    Image from @replicate's post
  3. Mark ZuckerbergXAI score5

    Meta's Muse app lets users create their own Muse today

    AIMark Zuckerberg announced that the Muse app is available now, and users can start making their own Muse at muse.ai. The post gives no details on what a Muse is or what the app does beyond creation.

  4. World LabsOfficialAI score32

    World Labs' Atlas keeps 3D consistency using spatial context

    AIWorld Labs says its Atlas model maintains 3D consistency during generation by building a spatial context from inputs and previously generated frames. The company adds that real-world images can be added to this spatial context to explore reconstructions of real-world scenes.

    Video from @theworldlabs's post

Sep 4

Sep 4Fri
  1. Mustafa SuleymanXAI score25

    Microsoft makes MAI image model 2.6 Flash available in Foundry

    AIMicrosoft's MAI-Image-2.6-Flash is now available in Microsoft Foundry and the MAI Playground. The post links to a Foundry model card for details, but gives no further specifications, pricing, or benchmarks.

  2. Mustafa SuleymanXAI score21

    MAI-Image-2.6-Flash generates images twice as fast as GPT-Image-2

    AIMicrosoft's MAI-Image-2.6-Flash generates images 2x faster than GPT-Image-2, which the post calls the best model in the world. It also uses 72% less GPU, enabling a lower price and what the post describes as the best price-performance score.

    Image from @mustafasuleyman's post
  3. Google AI StudioOfficialAI score44

    Google's Lyria 3.5 music model now available in AI Studio and Gemini

    AIGoogle has made Lyria 3.5, its best-sounding music generation model, available in AI Studio, through the Gemini API, and in the Gemini app. The model produces more expressive vocals and richer musical arrangements, enabling higher-fidelity tracks.

    Video from @GoogleAIStudio's post

Sep 3

Sep 3Thu
  1. Midjourney UpdatesOfficialAI score52

    Midjourney's alpha adds v8.2 edit model with lightbox editor

    AIMidjourney's alpha site now runs the new v8.2 edit model, with an editor built into the lightbox. Users can edit images with plain-text instructions, attach up to 4 reference images, and view all session edits in one place. The update also adds an early Change Style feature, and the team says speed and error messaging have improved, while drag and drop and the prompt bar are still in progress.

  2. RadixArkOfficialAI score16

    RadixArk's Miles supports shared post-training for VLMs and diffusion models

    AIRadixArk says its latest blog shows how Miles supports multimodal learning through a shared post-training design for vision-language models and diffusion models. The post frames cross-modal learning as necessary for AI to understand and recreate the real world.

    Image from @radixark's post
  3. Awni HannunXAI score51

    Mirai releases speculative decoding in Uzu for Qwen3.6-27B on Apple M5 Max

    AIMirai is releasing speculative decoding in its Uzu inference engine, starting with Qwen3.6-27B. The quoted post reports 105 output tokens per second on an Apple M5 Max with 128 GB of unified memory, 2.9× faster than the fastest MLX speculative-decoding implementation Mirai benchmarked. The stack combines DFlash with Mirai's Weaver model, tree-based speculative decoding, Mirai quantization, and Metal kernels for Apple silicon.

Sep 1

Sep 1Tue
  1. World LabsOfficialAI score35

    World Labs' Atlas generates precisely controlled 1440p video from reference images

    AIWorld Labs says its Atlas model places multiple input views within a spatial context to generate image and video frames with precise control. The company demonstrates a hand-designed camera trajectory that produces a one-minute video at 1440p resolution from seven reference images.

    Video from @theworldlabs's post
  2. World LabsOfficialAI score46

    World Labs unveils Atlas, a multimodal world model with 3D reconstruction

    AIWorld Labs has introduced Atlas, which it describes as the first multimodal world model that generates image and video frames with pixel-perfect camera control. The model can also reconstruct the generated frames in 3D, letting users move the camera and simulate space and time.

    Video from @theworldlabs's post
  3. Bryan CatanzaroXAI score13

    NVIDIA's DLSS 5 Neural Rendering Advances Real-Time Graphics

    AIBryan Catanzaro, NVIDIA's VP of applied deep learning research, called DLSS 5 Neural Rendering a major step toward the future of real-time graphics and the culmination of ten years of research. He described the technology as controllable, consistent, and fast, and said he looks forward to seeing what game developers build with it.

  4. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score60

    Shanghai AI Lab releases Intern Lumina U2 unified multimodal model on Hugging Face

    AIShanghai AI Lab's InternLM has published Intern Lumina U2, a 16B-parameter MoE model with 1B active parameters that handles text QA, image generation and editing, and image, video, and 3D understanding. The model uses an 8-codebook fully-discrete visual representation built on AToken. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs under Apache 2.0, with the technical report still listed as coming soon.

    Why it matters: The model unifies text, image, video, and 3D understanding with image generation in one framework, a broader scope than single-modality releases.

Aug 28

Aug 28Fri

Aug 27

Aug 27Thu
  1. RadixArkOfficialAI score12

    RadixArk publishes LoRA SFT guide for diffusion models

    AIRadixArk, a publisher on X, shares a full guide on LoRA supervised fine-tuning for the H3 model in its miles_diffusion repository on GitHub. The post itself contains only a link to the documentation, so no further details about the method or results are given.

  2. RadixArkOfficialAI score34

    RadixArk adds LoRA SFT to Miles-diffusion for targeted post-training

    AIRadixArk introduced LoRA SFT in Miles-diffusion for fast, targeted post-training of diffusion models. The company trained a rank-64 LoRA adapter for MiniMax H3 to improve physical realism, using 254 curated training windows and under 3 hours on 8 GPUs. The adapter can be exported to safetensors and served directly with SGLang without retraining the full model.

    Image from @radixark's post

Aug 26

Aug 26Wed
  1. Bryan CatanzaroXAI score46

    NVIDIA Releases DLSS 4.5 Ray Reconstruction with Better Image Quality

    AINVIDIA's DLSS 4.5 Ray Reconstruction is now available, using a second-generation joint denoiser and super-resolution model. According to the post, it delivers much better image quality at the same compute cost, pushing the trade-off between image quality and rendering cost further.