Skip to contentSkip to stories

Updated

#Image generation

Showing low-relevance items too. Hide low-relevance items

Jun 9

Jun 9Tue
  1. ByteDance · new models on Hugging FaceOfficialAI score28

    ByteDance releases Sa2VA-Qwen3-VL-4B-SAM3 for image and video referring segmentation

    AIByteDance's Sa2VA-Qwen3-VL-4B-SAM3 is built on Qwen3-VL-4B-Instruct with a SAM3 grounding encoder and produces dense image and video referring segmentation alongside chat. It reports 83.7 cIoU on RefCOCO val, 65.3 J&F on MeViS (val_u), and 77.1 on Ref-DAVIS17. The checkpoint is self-contained and loads on Hugging Face with trust_remote_code=True, with no extra packages required.

  2. Stability AIOfficialAI score23

    Stable Audio 3.0 powers creative stem remixing exploration

    AIStability AI says Stable Audio 3.0 was built to support exploratory audio work such as stem remixing. A quoted post from @teropa reports using the Medium model with an init_audio input and an init_noise_level of 0.4–0.5, with empty prompts.

Jun 8

Jun 8Mon
  1. ByteDance · new models on Hugging FaceOfficialAI score46

    ByteDance Open-Sources Bernini-R 1.3B Video Diffusion Renderer on Hugging Face

    AIByteDance has open-sourced the 1.3B-parameter weights of its Bernini Renderer (Bernini-R), available on Hugging Face as ByteDance/Bernini-R-1.3B-Diffusers. Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, but lags on complex tasks such as human generation. The release requires a CUDA GPU, with an H100 recommended for FlashAttention-3.

Jun 2

Jun 2Tue
  1. ByteDance · new models on Hugging FaceOfficialAI score44

    ByteDance Releases Bernini-R Diffusers Weights for Video Generation and Editing

    AIByteDance has open-sourced the inference code and model weights of the Bernini Renderer (Bernini-R), a DiT-based renderer paired with an MLLM-based semantic planner for video generation and editing. A diffusers-format version, ByteDance/Bernini-R-Diffusers, bundles the Wan2.2 base components with the Bernini-R transformer weights for direct loading, and the framework requires a CUDA GPU with PyTorch 2.5.1+cu124.

Jun 1

Jun 1Mon
  1. PaddlePaddleOfficialAI score36

    PaddleOCR and ERNIE Image now available as official Dify plugins

    AIPaddleOCR and ERNIE Image are now available as official Dify plugins, bringing document parsing and image generation into Dify's agent workflows. PaddleOCR, powered by PP-OCRv5, PP-StructureV3, and PaddleOCR-VL, turns images, scanned PDFs, and multilingual documents into structured data for chunking, vectorization, and RAG, with private or on-prem deployment supported. ERNIE Image offers free generation, a Turbo mode with 8-step inference, and an OpenAI-style API.

    Image from @PaddlePaddle's post

May 29

May 29Fri
  1. Fei-Fei LiXAI score38

    Fei-Fei Li Highlights GPIC, a Permissive Image Corpus for Visual Generation

    AIFei-Fei Li praised GPIC, a new benchmark dataset for visual generation built for modern large-scale generative models. The corpus includes 100M VLM-captioned image-text pairs for training and 1M pairs for benchmarking, totaling about 28 trillion pixels. It is centrally hosted and fully permissive for research and commercial use.

May 20

May 20Wed

Apr 21

Apr 21Tue
  1. Nick TurleyXAI score67

    ChatGPT Images 2.0 launches with better instruction following and dense text rendering

    AINick Turley announced ChatGPT Images 2.0 as a major advance in image generation, citing better adherence to detailed instructions, rendering of dense text, and more accurate understanding of the world. He said the model can spend extra time planning and refining outputs for tasks needing more accuracy and clarity, and that users have generated over 1 billion images with ChatGPT.

    Why it matters: The post names concrete gains in instruction following, dense text rendering, and optional extended thinking for image output, which helps readers gauge practical scope.

    Image from @nickaturley's post
  2. NVIDIA AI DeveloperOfficialAI score39

    RL post-training offers a steerable alternative to CFG for image generation

    AIResearchers introduced a simple, sample-efficient online reinforcement learning technique for post-training image generation models. It is presented as a possible steerable alternative to classifier-free guidance (CFG) that can be driven by any scalar reward, including human preference.

Apr 8

Apr 8Wed
  1. Stability AIOfficialAI score44

    Stability AI launches Brand Studio, a creative production platform built around brand identity

    AIStability AI has introduced Brand Studio, an end-to-end creative production platform for enterprise teams that builds around each brand's identity. Its Brand Central hub supports custom Brand ID models and Campaigns, while Producer Mode turns prompts into step-by-step production plans. Curated Model Routing selects models including Stable Diffusion, Nano Banana, and Seedream, and new Precision Inpainting and Product Insertion tools enable targeted edits.

Apr 6

Apr 6Mon
  1. Black Forest Labs · new models on Hugging FaceOfficialAI score41

    FLUX.2 Small Decoder offers faster, lower-VRAM drop-in replacement for FLUX.2 decoder

    AIBlack Forest Labs released FLUX.2 Small Decoder, a distilled VAE decoder that works as a drop-in replacement for the standard FLUX.2 decoder on Hugging Face. It decodes about 1.4x faster and uses about 1.4x less VRAM at decode time, with ~28M decoder parameters versus ~50M in the full decoder and minimal quality loss. It is available under the Apache 2.0 license and is compatible with FLUX.2-klein-4B, FLUX.2-klein-9B, FLUX.2-klein-9b-kv, and FLUX.2-dev.

Mar 12

Mar 12Thu
  1. Intern Large ModelsOfficialAI score47

    InternVL-U: Open-Source 4B Unified Model for Reasoning, Generation, and Editing

    AIInternVL-U is a lightweight 4B unified multimodal model that combines reasoning, generation, and editing in one framework, according to Intern Large Models. The post says it uses unified contextual modeling, modality-specific modular design, and decoupled visual representations to balance performance and efficiency. It reportedly outperforms unified baselines more than 3× its size on text rendering, scientific reasoning, and spatially grounded generation and editing, and is open-source on GitHub and Hugging Face.

    Image from @intern_lm's post

Mar 11

Mar 11Wed
  1. Nano Banana 2.1OfficialAI score62

    How to get the most out of Nano Banana 2 for image generation

    AINano Banana 2, also called Gemini 3.1 Flash Image, adds visual grounding with Google Search, 512px resolutions, and extreme aspect ratios of 1:8 and 1:4. The guide advises using it as the default for new projects, with Nano Banana Pro reserved for complex prompts it fails, and keeping Thinking mode off by default.

Mar 9

Mar 9Mon
  1. Black Forest Labs · new models on Hugging FaceOfficialAI score39

    Black Forest Labs releases FLUX.2 [klein] 9B-KV with KV-cache for faster multi-reference editing

    AIBlack Forest Labs has released FLUX.2 [klein] 9B-KV, a variant of FLUX.2 [klein] 9B that caches reference-image key-value pairs to speed up multi-reference editing by up to 2.5 times. The 9B flow model, which uses an 8B Qwen3 text embedder and is step-distilled to 4 inference steps, is available for non-commercial use under the FLUX Non-Commercial License and fits in about 29GB VRAM.

Feb 28

Feb 28Sat
  1. Nano Banana 2.1OfficialAI score18

    Nano Banana 2 can turn any subject into an enamel pin

    AIGoogle's Nano Banana account says users can turn anything into an enamel pin with Nano Banana 2, either by selecting the "Enamel pin" style in the Gemini app or by using a prompt. The example prompt asks for just the subject as a gold enamel pin shown as a minimal photo on a desk.

    Image from @NanoBanana's post

Feb 27

Feb 27Fri
  1. Nano Banana 2.1OfficialAI score10

    Nano Banana 2 generates cute spherical animal characters from a prompt

    AIGoogle's Nano Banana 2 is presented as the best model for a prompt that asks for a cute, nearly spherical animal with oversized eyes and no separate head, shown in its natural habitat. The example prompt specifies cartoonishly cute features with the eyes looking up cheekily.

    Image from @NanoBanana's post

Feb 26

Feb 26Thu
  1. Oriol VinyalsXAI score60

    Nano Banana 2 debuts at #1 in Image Arena text-to-image ranking

    AINano Banana 2, officially released as Gemini 3.1 Flash Image Preview, ranks first in Image Arena text-to-image with a score of 1279. The quoted post says it also ties for first in single-image editing at 1407 and costs $0.067 per image, about half the price of Nano Banana Pro.

    Image from @OriolVinyalsML's post
  2. Nano Banana 2.1OfficialAI score67

    Google introduces Nano Banana 2, its best image generation and editing model

    AINano Banana announces Nano Banana 2, which it describes as its best image generation and editing model yet. The model can be tried in the Gemini app, Google AI Studio, and other places the post does not specify.

    Why it matters: The post names the access points for Nano Banana 2, which helps readers see where the image generation and editing model can be tried.

Feb 24

Feb 24Tue
  1. Nano Banana 2.1OfficialAI score10

    Nano Banana prompt turns an image into homemade cookies on a plate

    AINano Banana, the account linked to Google and Gemini, shares a prompt asking an image model to imagine the image as a simple cookie cutter and show the resulting unfrosted homemade cookies on a plate without showing the cutter. The post is a creative prompt example rather than an announcement of a new feature.

    Image from @NanoBanana's post

Feb 18

Feb 18Wed
  1. Nano Banana 2.1OfficialAI score6

    Nano Banana shares a prompt for a simple 9:16 gradient wallpaper

    AINano Banana, Google's Gemini image account, shares a prompt for a 9:16 wallpaper with negative space near the top and bottom, a subtle top-to-bottom gradient, no blur, and no cropping. The post advises keeping the scene simple and placing the subject and style at the end of the prompt, with example images attached.

    Image from @NanoBanana's post

Feb 16

Feb 16Mon
  1. Nano Banana 2.1OfficialAI score14

    Nano Banana shares a highly detailed prompt example for image generation

    AINano Banana (@NanoBanana) posts a sample prompt showing how specific image-generation instructions can be, describing a San Francisco cafe scene with a couple by the window, a sleeping black-and-white French bulldog, and a mirrored "Nano's Cafe" gold lettering on the glass. The post offers the detailed prompt as an illustration of how far users can go with specificity, covering clothing, props, lighting, and poses.

    Image from @NanoBanana's post

Jan 15

Jan 15Thu
  1. Nano Banana 2.1OfficialAI score18

    Risograph prompt yields vibrant icon grid of cute, weird real animals

    AINano Banana (@NanoBanana) suggests adding the word "risograph" to image prompts for a distinctive print effect. The example prompt asks for a 2x2 grid of icons on pure white with thick black outlines, stochastic stippling, and vibrant, unfaded color, here themed "cute but weird animals (that are real)."

    Image from @NanoBanana's post

Jan 14

Jan 14Wed
  1. Black Forest Labs · new models on Hugging FaceOfficialAI score62

    Black Forest Labs releases FLUX.2 [klein] 4B image model under Apache 2.0

    AIBlack Forest Labs released FLUX.2 [klein] 4B, a 4 billion parameter model that unifies text-to-image generation and image editing with multi-reference support. The source says it runs on consumer GPUs such as the RTX 3090 or 4070 with about 13GB VRAM, and its open weights are available under the Apache 2.0 license.

    Why it matters: The source specifies a 4 billion parameter model running on about 13GB VRAM under Apache 2.0, which helps readers judge whether local image generation fits their hardware.

  2. Black Forest Labs · new models on Hugging FaceOfficialAI score54

    Black Forest Labs releases FLUX.2 [klein] 4B Base on Hugging Face

    AIBlack Forest Labs has published FLUX.2 [klein] 4B Base, a 4 billion parameter text-to-image model that also supports multi-reference editing. The model is undistilled, is released with open weights under Apache 2.0, and is described as fitting in about 13GB VRAM on cards such as the RTX 3090 or 4070, with reference code available in its GitHub repository and support in ComfyUI and Diffusers.

  3. Black Forest Labs · new models on Hugging FaceOfficialAI score46

    FLUX.2 [klein] 9B Base Released on Hugging Face as Undistilled Open-Weight Model

    AIBlack Forest Labs has released FLUX.2 [klein] 9B Base, a 9 billion parameter undistilled rectified flow transformer with open weights for text-to-image generation and multi-reference editing. The model is intended for fine-tuning, LoRA training, and research, and fits in about 29GB VRAM on NVIDIA RTX 4090-class GPUs. A reference implementation is available on GitHub, and the model works with ComfyUI and Diffusers.

Jan 13

Jan 13Tue
  1. Z.ai Release NotesOfficialAI score34

    GLM-Image launches as Z.ai's image generator trained on domestic chips

    AIZ.ai has launched GLM-Image, an image generation model built on a multimodal architecture that combines autoregressive semantic understanding with diffusion-based decoding. The update improves knowledge-intensive generation and makes text rendering inside images more stable and accurate, suiting commercial design and educational illustrations.

Jan 9

Jan 9Fri
  1. Nano Banana 2.1OfficialAI score12

    Nano Banana shares a prompt for 3x3 3D icon grids

    AINano Banana posts a prompt template for generating a 3x3 grid of colorful, tactile 3D icons on a white background with no text. The post lists example themes: dogs with different emotions, bananas, January, and the same cat in different emotions.

    Image from @NanoBanana's post
  2. BAAIOfficialAI score47

    DrugCLIP screens 10 trillion protein-molecule pairs per day for drug discovery

    AITsinghua AIR and BAAI's DrugCLIP screened 10,000 proteins against 500 million molecules, identifying over 2 million drug candidates. The post claims a 1-million-fold speedup, reaching 10 trillion protein-molecule pairs per day, and positions DrugCLIP as bridging AlphaFold structures to drug candidates. The work is published in Science, with a platform available at drugclip.com.

    Image from @BAAIBeijing's post

Dec 19, 2025

Dec 19, 2025Fri

Dec 17, 2025

Dec 17, 2025Wed
  1. Nano Banana 2.1OfficialAI score4

    Nano Banana shares a Christmas card prompt for 2025

    AINano Banana, Google's Gemini image account, posts a sample prompt for a traditional Christmas card featuring a couple in Christmas hats. The prompt specifies the text "Happy Christmas 2025" and a classical, friendly design in a winter wonderland setting.

    Image from @NanoBanana's post

Dec 16, 2025

Dec 16, 2025Tue

Dec 12, 2025

Dec 12, 2025Fri
  1. Nano Banana 2.1OfficialAI score14

    Nano Banana Pro creates isometric 3D miniature room scenes from photos

    AINano Banana, the account owned by Google and Gemini, thanks @Arminn_Ai for inspiring a post showing 3D miniature isometric room images made with Nano Banana Pro. The post says users can upload their own photo to generate a personalized chibi-style figure in a cube-shaped room and offers to share the full prompt on request.

  2. Nano Banana 2.1OfficialAI score24

    Nano Banana shares a prompt for cute isometric diorama lamps

    AINano Banana, the account of Google's Gemini team, shares a prompt template for generating cute isometric 3D cube diorama lamps. The prompt specifies internal lighting, chibi figurine styling, matte PVC material, a neutral background in a dark room, many small details, and subtle dust and scratch textures.

    Image from @NanoBanana's post
  3. Apple · new models on Hugging FaceOfficialAI score46

    Apple's SHARP Turns a Single Photo into a 3D Scene in Under a Second

    AIApple has released SHARP, a model that generates a 3D Gaussian representation of a scene from a single photograph in less than a second on a standard GPU. The output renders in real time as high-resolution photorealistic views of nearby camera positions, with metric absolute scale, and the paper reports reductions of 25–34% in LPIPS and 21–43% in DISTS versus the best prior model.

Dec 11, 2025

Dec 11, 2025Thu