Grok Imagine Image 2.0 is xAI's image generator and editor: 1k or 2k output, low or medium quality, thirteen aspect ratios, and up to five reference images per edit.


Proofread against the brief
6 of 6 requested lines spelled correctly on the first run. Select a line to find it in the image.
Risograph-style event poster for an invented public observatory night. Title in large bold condensed capitals across the top: "A NIGHT OF FALLING STARS". Directly beneath it a smaller line: "Perseid viewing on the west lawn". … Near the bottom a row of three short labels: "TELESCOPES", "STAR MAPS", "HOT COCOA". Footer line at the very bottom: "Saturday 12 August - 9 PM until late". Two-colour print in navy and warm orange on cream paper…
Four settings, one prompt
One brief — a brass compass with a 0–360 ring and the needle pointing north — at every setting. Select a cell to inspect the dial at native pixels.

The only cell where the needle points north. The ring still runs 340, 0, 100, 90, 100 — and a compass has no 380.
Our read: medium and 2k buy texture and pixels, not correctness. Add dense numbers in a design tool.
Change the words, keep the picture

“…the second line reads PLUM CAKE instead of LEMON TART… replace the lemon slice with a chalk plum. Keep TODAY, 4.50… exactly the same.” 1k medium, 26.8 s, 25 credits.
Same prompt, both versions. The first model also spelled every line correctly — the difference is in everything around the words.


Thirteen shapes
Highlighted shapes have a real output from our runs. Select one to view it at its true proportions.

Exactly as it should print: "HOT COCOA", not "a cocoa label".
"Across the top", "directly beneath", "at the very bottom".
"Large bold condensed capitals" versus "a smaller line".
"All text straight, evenly spaced and perfectly legible" closed both of our text prompts.
Short numbers like 4.50 or 35 minutes held. A full 0–360 scale did not, at any setting.
Title, subtitle, labels and date all held.
Edit the dish name as the menu changes.
Two columns and numbered steps, spelled right.
A name and one descriptor line.
Big type over an illustration.
One headline, one line, 1:1 or 9:16.
9:19.5 and 9:20 fill the screen.
2:1 or 20:9; use 2k for retina.
Fictional faces, film-still grade.
Change one word, keep the layout.
Structure yes; check every number.
Up to five references per edit.
You need motion. Generate the still here, then animate it with the Grok Imagine video generator.
You need long or technical copy. Dense numerals and small print need a model built for layout work, such as GPT Image 2, plus a human proofread.
You want to compare first. Text to image runs the same brief across every image model on Voor.
Grok Imagine Image 2.0 is xAI's second-generation image model. On Voor it generates from text and edits up to five reference images, at 1k or 2k, low or medium quality.
Yes. Both spellings refer to the same release; xAI versions it as 2.0.
In our tests Grok Imagine Image 2.0 spelled every requested line correctly on an event poster (6 of 6) and a recipe card (11 of 11). It failed on a compass dial: the 0–360 degree numbers came out repeated and out of order at every setting we tried.
New images: 14 credits at 1k low, 21 at 1k medium or 2k low, 28 at 2k medium. Edits add a little per reference: 25 credits for one image at 1k medium, up to 46 for five at 2k medium. New accounts get welcome credits.
Low finished in 17 to 23 seconds in our runs and is fine for checking a layout. Medium took 71 to 102 seconds for new images but rendered finer surface detail in our side-by-side, so switch to it once the composition is settled.
1k is about one megapixel (1024 × 1024 square); 2k is about four (2048 × 2048). Choose 2k for print or cropping.
Yes. We changed LEMON TART to PLUM CAKE on a chalkboard and rewrote the date line of a finished poster; everything else held. The edit keeps the first image's aspect ratio on auto, but follows the resolution setting, so a 2k poster edited at 1k comes back smaller.
On our poster prompt the original returned 864 × 1152 in 6 seconds for 7 credits, also spelled correctly. Grok Imagine Image 2.0 adds 2k output, richer layout, quality control and multi-image edits, at more cost and wait.
Thirteen, from 2:1 to 1:2, plus auto for edits, which keeps the first reference image's shape.
Start at 1k low to check the layout, then render the keeper at medium or 2k and proofread it at full size.
Create with Grok Imagine Image 2.0Optional cookies help us understand usage and measure ads. Essential cookies stay on so Voor works.