Skip to contentSkip to stories

Updated

#Image generation

Showing low-relevance items too. Hide low-relevance items

Sep 23

Sep 23Wed
  1. World LabsOfficialAI score27

    World Labs previews Chisel, a world-building feature in Atlas beta

    AIWorld Labs has previewed Chisel, a feature in its Atlas beta that lets users block out a world and have Atlas bring it to life. The post offers no further technical details, pricing, or availability information.

    Video from @theworldlabs's post
  2. Microsoft AIOfficialAI score29

    Microsoft's MAI-Image-2.6 now available on Foundry and OpenRouter

    AIMicrosoft AI announced that its MAI-Image-2.6 image model can be tried on Microsoft Foundry and OpenRouter. The post provides links to both platforms but no further details on capabilities, pricing, or benchmarks.

  3. Microsoft AIOfficialAI score22

    Microsoft AI's image model climbs 864 to 1147 Elo in nine months

    AIMicrosoft AI reports its image model rose from 864 to 1147 Elo in nine months, a gain of 377 points per year. The company says that pace is 1.6 times the next-fastest climb on the leaderboard. The post does not name the model.

    Video from @MicrosoftAI's post
  4. Google GeminiOfficialAI score10

    Gemini now generates and edits images through Picsart integration

    AIGoogle Gemini users can ask Gemini to generate multimedia, change backgrounds, and upscale image quality via Picsart. The post gives an example prompt: "Create a vibrant modern logo for an artisanal matcha café in Picsart."

    Image from @GeminiApp's post
  5. Google LabsOfficialAI score29

    Google Labs Releases Six Flow Tools Built by Creatives in Sound, Design, and Content

    AIGoogle Labs released six new Google Flow Tools built by creatives across architecture, sound design, and digital content, including Mondo Sónico, CaptionCast, ThumbnailForge, Surface, CollageMotion Pro, and SwissFlow Studio. Each tool targets a specific workflow, such as generating synchronized audio stems, transcribing and styling captions, or producing animated collages from text prompts. Users can try the tools, duplicate and remix them, or build their own by describing a task in Google Flow.

  6. Comfy BlogOfficialAI score62

    Comfy Router launches one API for frontier image, video, 3D, and audio models

    AIComfy Router is now live on the Comfy Developer Platform, giving developers one API to call frontier image, video, 3D, and audio models. Day one models include Seedance 2.5, MiniMax H3, Nano Banana Pro, GPT Image 2, Kling, and Black Forest Labs, and the provider for each job is selectable. Requests fail rather than silently switching providers, and inputs and outputs are deleted after 24 hours.

    Why it matters: The post shows how one API key and a provider parameter let developers swap routes for media models without rewriting calls, with failed requests reporting the provider.

  7. QwenOfficialAI score60

    Qwen Intelligence launches three mobile agents and opens its benchmark suite

    AIAlibaba's Qwen launched Qwen Intelligence with three mobile agents: a Mobile Planner Agent, a Mobile-Use Agent, and a Mobile Creative Agent. The post reports benchmark results including MobileWorld 82.1, MobileWorld-Real 92.2, and AndroidDaily 97.2, plus a 90% end-to-end success rate, and says the MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety benchmarks are open.

    Image from @Alibaba_Qwen's post
  8. KrASIA · Big TechNewsAI score46

    Tencent Hy Image 3.5 preview refined through its consumer and business products

    AITencent has released a preview of its Hy Image 3.5 image generation model, which product teams across Yuanbao, WorkRally, Ima, and other services are helping refine through co-design. Tencent Cloud prices the model at USD 0.024 per 2K output image, and it supports text-to-image and image-to-image generation with up to five reference images. Tencent said an internal blind evaluation found it on par with ByteDance's Seedream 5.0 Pro and slightly better than Nano-Banana Pro and Qwen-Image-3.0 Pro.

  9. Tencent HyOfficialAI score22

    Tencent Hunyuan's Hy Image3.5 preview free for two weeks on OnSolo

    AITencent Hunyuan has made its Hy Image3.5 preview available on the OnSolo platform, free for two weeks. The model targets short drama character sheets, full-motion video game assets, and keyframes, with characters kept consistent across episodes and edits that refine rather than regenerate images. OnSolo's background post says the preview supports 5 references at 2K resolution and is free for Members during the two-week window.

Sep 22

Sep 22Tue
  1. Tencent HyOfficialAI score43

    Tencent Hunyuan previews Hy Image3.5 in ComfyUI

    AITencent Hunyuan's Hy Image3.5 preview is now available in ComfyUI, with a claimed 30% higher win rate in human evaluation than Hy Image3.0. The model handles text-to-image and image-to-image in one model at up to 2K resolution, with multilingual text and small print rendering correctly. It also keeps identity and product features consistent across scene, outfit, and style changes.

  2. ModelScopeOfficialAI score62

    inclusionAI open-sources Ming-Image-0.1-Design models for visual design

    AIinclusionAI open-sources the Ming-Image-0.1-Design family, two complementary 6B models for visual-design workflows, under an MIT License. Design generates complete UIs, dashboards, infographics, and posters up to 2048×2048 with native transparent RGBA output, and Layer decomposes flattened graphics into independently editable RGBA layers.

    Image from @ModelScope2022's post
  3. Daniel HanXAI score22

    Unsloth Desktop hotfix adds Qwen-Image-2.1 image editing and fixes

    AIUnsloth Desktop received a hotfix update adding image editing for Qwen-Image-2.1. The update also fixes diffusers update issues, GGUF loading failures for Qwen-Image, and black artifacts during diffusion on A100 and consumer GPUs. Users should receive a banner prompting them to update.

  4. Unsloth AIOfficialAI score26

    Qwen-Image-2.1 FP8 and GGUF quants now run in Unsloth Desktop

    AIUnsloth announced that Qwen-Image-2.1 FP8 and GGUF quantized versions should now run properly in Unsloth Desktop. The app supports both image generation and image editing with these quants. Further details are available on the Unsloth GitHub repository.

    Image from @UnslothAI's post
  5. Alex AlbertXAI score18

    Opus 5.5 builds a 1906 San Francisco street scene in Blender

    AIAlex Albert says Opus 5.5 has improved 3D modeling and vision for Blender work, letting users build an entire world from a single prompt. He shares a historically accurate render of San Francisco's Market Street in 1906, before the earthquake.

    Video from @alexalbert__'s post
  6. Felix RiesebergXAI score47

    Anthropic's Opus 5.5 praised for natural writing and computer art

    AIAnthropic's Felix Rieseberg says the new Opus 5.5 model writes more naturally than earlier models. He also highlights its strong generative computer art and drawing ability, noting it is not an image model yet produces attractive visuals.

  7. Daniel HanXAI score42

    Qwen-Image-2.1 runs locally in Unsloth Desktop via INT8, FP8, GGUF

    AIDaniel Han says Qwen-Image-2.1 works in Unsloth Desktop through INT8, FP8, and GGUF builds, with Unsloth also releasing dynamic GGUFs for it. Pinned RAM offloading lets INT8 and FP8 fit under 6–8GB of VRAM while remaining relatively fast. The linked Unsloth post says the 7B model runs on 12GB VRAM and performs on par with Nano Banana 2.0.

  8. Unsloth AIOfficialAI score70

    Qwen-Image-2.1 runs locally on 12GB VRAM using Unsloth GGUFs

    AIUnsloth says the 7B Qwen-Image-2.1 text-to-image and editing model can run locally on 12GB VRAM using its GGUF builds. It also states that the model performs on par with Nano Banana 2.0, and that Dynamic FP8 can run on 6GB of VRAM via offloading for higher quality. The image lists int8 at 7.26 GB with mean LPIPS 0.064 and fp8 at 7.12 GB with mean LPIPS 0.112, and says int8 is the default.

    Why it matters: The post gives concrete local-run settings, VRAM figures, and GGUF and FP8 options, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  9. Tencent HyOfficialAI score34

    Tencent Hunyuan's Hy Image3.5 preview launches free on Miora for two weeks

    AITencent Hunyuan has released a preview of Hy Image3.5 on Miora, a design platform, and is offering it free for two weeks. The model keeps existing canvas workflows and remembers users' brand rules while making edits. Miora's background post says the free period runs through October 7 and supports up to 2K output for text-to-image and image-to-image.

  10. Tencent HyOfficialAI score38

    Tencent Hunyuan shares showcase cases from Hy Image3.5 preview

    AITencent Hunyuan posted a set of showcase images generated by its Hy Image3.5 preview model. The post presents the outputs as examples of the model's capabilities, without accompanying detail on methods or benchmarks.

    Image from @TencentHunyuan's post
  11. Tencent HyOfficialAI score8

    Tencent Hunyuan confirms its model can edit images

    AITencent Hunyuan replied that its model can edit images, not just generate them, answering a user's question. The post names no specific model, version, or benchmark.

Sep 21

Sep 21Mon
  1. Tencent HyOfficialAI score38

    Tencent Hunyuan releases Hy Image3.5 preview for image generation

    AITencent Hunyuan has launched a preview of Hy Image3.5, which it says wins 30% more often than Hy Image3.0 in human evaluation. The model supports text-to-image and image-to-image generation at up to 2K resolution with improved consistency, and is priced at $0.024 per image on the Tencent Cloud API, with reference images free. Two weeks of free access is offered through OnSolo and Miora.

    Video from @TencentHunyuan's post
  2. Xiaomi MiMoOfficialAI score36

    Xiaomi MiMo-V2.6 unifies code, design, and tool use across creative outputs

    AIXiaomi's MiMo-V2.6 combines code, design, and tool use to build frontend interfaces, presentations, Figma-linked visual assets, and videos. The post says MiMo-V2.5-TTS supports narration in video production, and that the model can compose music, including an orchestral piece for around ten instruments that can be converted to MIDI. On Design Arena, the Pro version reportedly performs comparably to Claude Opus 5 and GPT-5.6 Sol.

    Image from @XiaomiMiMo's post
  3. Xiaomi MiMoOfficialAI score44

    MiMo-V2.6 builds and interacts with 3D worlds from text, images, or video

    AIXiaomi's MiMo-V2.6 combines 3D spatial reasoning, multimodal perception, and computer use to turn text, images, or video into playable 3D worlds. The model coordinates agents to build scenes, write interaction logic, and refine results, and can create Blender objects for animation, 3D printing, and games. It also controls a Franka Panda arm in simulation via visual feedback and uses desktop tools to process data, inspecting results to adjust its next actions.

    Video from @XiaomiMiMo's post
  4. Xiaomi MiMoOfficialAI score34

    Xiaomi's MiMo Gallery showcases outputs generated entirely by MiMo-V2.6

    AIXiaomi's MiMo team released MiMo Gallery, a showcase where every 3D model, game, slide deck, video, music piece, and image was generated by MiMo-V2.6. The post previews MiMo-V2.6 as an upcoming model and points readers to the gallery for a first look.

    Video from @XiaomiMiMo's post

Sep 20

Sep 20Sun
  1. QwenOfficialAI score38

    Qwen-Image-2.1 launches live Spaces demo for generation and editing

    AIQwen-Image-2.1 is now available as a live Hugging Face Spaces demo, letting users try image generation and editing in the browser without setup. The demo runs on a single checkpoint that handles both tasks. Background notes describe a 7B-parameter model supporting up to 10 image references, with integration into diffusers and ComfyUI.

  2. QwenOfficialAI score34

    Qwen-Image-2.1 now supported in ComfyUI for image generation

    AIQwen-Image-2.1 is now supported in ComfyUI, and Qwen invites users to try it and share their creations. ComfyUI describes it as an open-weights 7B checkpoint that handles both generation and editing, with native 2K image generation and instruction editing from up to 10 reference images in one pass.

  3. QwenOfficialAI score34

    Qwen-Image-2.1 edits three marked regions in one prompt

    AIQwen-Image-2.1 supports several ways to specify local edits, including circles marking regions on an image. In the example, the model removes a metal watch, changes hair color to black, and replaces clothing in three circled areas in a single pass.

    Image from @Alibaba_Qwen's post
  4. QwenOfficialAI score56

    Qwen-Image-2.1 releases open weights for image generation and editing

    AIAlibaba's Qwen team released Qwen-Image-2.1 as an open-weights image model for both generation and editing, with a lightweight 7B architecture. The model natively generates and edits RGBA layers, supports up to 10 reference images for editing, and is available on GitHub, ModelScope, and Hugging Face.

    Image from @Alibaba_Qwen's post