How to Make AI Images From Text Prompts
AI image generation has a low floor and a high ceiling. Getting a decent image requires nothing more than typing a sentence. Getting a great image consistently requires understanding how prompts translate into visual output, what each platform does well, and how to iterate efficiently. The steps below cover both the basics and the techniques that separate casual users from people who produce genuinely impressive results.
Step 1: Choose an AI Image Generator
Your choice of platform determines the quality ceiling, the cost per image, the interface complexity, and what types of images you can create. For beginners, the decision comes down to three practical factors: budget, technical comfort, and intended use.
If you want to start generating immediately without spending money, go to Leonardo AI (leonardo.ai) and create a free account. You get 150 tokens per day, enough for 5 to 10 images, with no credit card required and no trial period. The web interface is clean and intuitive, and the output quality is competitive with paid platforms. This is the recommended starting point for anyone who has never used an AI image generator before.
If you already pay for ChatGPT Plus ($20 per month), you have access to DALL-E through the ChatGPT interface. Simply ask ChatGPT to generate an image by describing what you want in conversation. This is the easiest workflow available because you can iterate through natural language without learning any new interface.
If you want the highest aesthetic quality and are willing to pay for it, subscribe to Midjourney at midjourney.com starting at $10 per month. Midjourney produces the most visually striking default output of any generator, with cinematic lighting and composition that makes even simple prompts produce impressive results.
If you want free, unlimited generation and are comfortable with technical setup, install Flux or Stable Diffusion locally. This requires a computer with an NVIDIA GPU (12GB VRAM for Flux, 8GB for Stable Diffusion) and approximately 30 to 60 minutes of initial configuration. Once set up, you generate as many images as you want with no ongoing costs.
Step 2: Write Your First Prompt
A prompt is the text description that tells the AI what to generate. The most important principle is specificity: the more precisely you describe what you want, the closer the output will match your vision. A prompt has several components that each influence the result.
Start with the subject. What is the main focus of the image? Be concrete: "a tabby cat with green eyes" is better than "a cat." Add context and setting: "sitting on a windowsill overlooking a rainy street at night." Include the medium or style: "photograph," "oil painting," "digital illustration," "3D render," "watercolor." Each style keyword activates different visual patterns in the model.
Add lighting description. This single element has more impact on the quality and mood of your image than almost any other prompt component. "Golden hour sunlight streaming through the window" versus "cool blue moonlight" versus "warm interior lamplight" will each produce dramatically different results from the same subject description.
Include technical references if you want photorealistic output. Camera and lens specifications like "shot on Sony A7R IV, 85mm f/1.4, shallow depth of field" tell the model to produce imagery with the visual characteristics associated with that equipment. Film stock references like "Kodak Portra 400" or "Fuji Velvia" add specific color science and grain characteristics.
Example beginner prompt: "A golden retriever puppy sitting in a field of wildflowers, photograph, golden hour sunlight, shallow depth of field, warm color palette"
Example advanced prompt: "A golden retriever puppy sitting in a meadow of purple lupines and orange poppies, late afternoon golden hour, warm sunlight backlighting the puppy's fur, shot on Canon EOS R5 with 85mm f/1.2 lens, shallow depth of field with creamy bokeh, Kodak Portra 400 color profile, professional pet photography"
Step 3: Set Generation Parameters
Before hitting generate, check and adjust the available settings. The specific options vary by platform, but the most common parameters include the following.
Aspect ratio controls the shape of your output image. Square (1:1) is the default on most platforms and works for general-purpose images. Landscape (16:9 or 3:2) works for desktop wallpapers, presentations, and cinematic compositions. Portrait (9:16 or 2:3) works for phone wallpapers, Pinterest pins, and vertical social media content. Match the aspect ratio to your intended use before generating rather than cropping afterward.
Number of images per generation is usually set to 4 by default. Generating multiple variations from the same prompt gives you options to choose from and helps you understand how the model interprets your description. Keep this at 4 for exploration, reduce to 1 or 2 once you have refined your prompt and want to save credits.
Model selection, available on platforms like Leonardo AI and Stable Diffusion interfaces, lets you choose between different base models optimized for different styles. Photorealistic models produce camera-like output. Anime models produce Japanese illustration styles. Artistic models produce painterly or illustrative output. Choose the model that matches your target aesthetic.
Quality or step count settings control how many refinement passes the model makes. Higher values produce more detailed, refined images but take longer and consume more credits. Start with the default setting, increase it when you need maximum quality for a specific image, and decrease it when you are exploring ideas quickly and quality is less critical.
Step 4: Generate and Evaluate the Output
Hit the generate button and wait for results. Generation typically takes 5 to 30 seconds depending on the platform, model, and quality settings. When your images appear, evaluate them against your vision rather than just picking the "best looking" one.
Check subject accuracy. Does the image contain what you described? Are the colors correct? Is the composition what you intended? If the subject is fundamentally wrong, your prompt needs restructuring rather than minor tweaks.
Check for anatomical issues. Examine hands, fingers, eyes, and teeth closely if the image contains people or animals. These are the areas where AI generators most frequently produce errors. Modern models handle anatomy much better than earlier versions, but complex poses still occasionally produce artifacts.
Check for coherence. Does the lighting make physical sense? Are shadows consistent? Do reflections appear where they should? In complex scenes, AI generators sometimes produce lighting that looks appealing at a glance but is physically impossible on closer inspection.
Check text if applicable. If your image includes any text elements (signs, labels, writing), verify that they are spelled correctly and legible. If they are not, you will likely need Ideogram for text-heavy images or plan to add text manually in post-processing.
Do not expect perfection on the first generation. AI image generation is an iterative process. Your first image gives you information about how the model interprets your prompt, which you use to refine in the next step.
Step 5: Refine Through Iteration
Iteration is where good AI images become great ones. You have several refinement techniques available depending on your platform.
Prompt adjustment is the most straightforward approach. If the image is close but not quite right, modify your prompt to address the specific issues. If the colors are too cool, add "warm color palette" or "golden tones." If the composition is too tight, add "wide shot" or "environmental portrait." If the style is too photographic, change "photograph" to "digital painting" or "concept art." Make one change at a time so you can identify which adjustments produce which effects.
Variation features let you generate multiple similar versions of an image you like without changing your prompt. Most platforms offer a "vary" or "variations" button that takes an existing generation and creates modified versions that preserve the overall composition while changing details. This is faster than re-prompting when you have an image that is 80% right.
Seed locking is available on most platforms (Midjourney, Leonardo, ComfyUI for Flux and Stable Diffusion). Every generated image has a seed number, a starting noise pattern that determines the composition. By locking the seed and changing only the prompt, you can modify specific elements of an image while maintaining the overall structure. This is essential for precise refinement.
Image-to-image generation lets you upload your best output as a reference for a new generation. The model uses the reference's composition and structure while applying your updated text prompt. This is the most powerful refinement technique for maintaining consistency while making significant changes to style, lighting, or specific elements.
On ChatGPT with DALL-E, you can simply describe what you want changed: "make the background darker," "remove the person on the right," "change the cat's color to orange." The conversational interface handles the technical details of re-generation and editing.
Step 6: Upscale and Finalize
Most AI image generators output at 1024x1024 pixels natively. This is sufficient for web use, social media, and screen display, but too small for print, posters, or large-format applications. Upscaling increases the resolution while adding plausible detail that the original did not contain.
Built-in upscalers are available on Midjourney (the "Upscale" button doubles or quadruples resolution), Leonardo AI (upscale feature on paid plans), and most other platforms. These are the easiest option because they work within the generation interface without requiring additional tools.
Dedicated AI upscalers like Topaz Photo AI, Magnific, and free options like Real-ESRGAN produce higher quality results than built-in upscalers in most cases. If you regularly need high-resolution output, investing in a dedicated upscaler improves quality noticeably compared to the upscale features built into generation platforms.
Post-processing in a photo editor is optional but common for professional use. Minor adjustments like cropping, color correction, sharpening, and adding text are easier and more precise in a photo editor than through prompt refinement. Adobe Photoshop, Canva, and GIMP all work for basic post-processing. If you need to composite AI-generated elements with real photographs, Photoshop's Generative Fill (powered by Adobe Firefly) creates seamless transitions between real and generated content.
Save your final image in the appropriate format: PNG for images that need transparency or lossless quality, JPEG for photographs and web use, and WebP for optimized web delivery. If you plan to revisit or modify the image later, also save the prompt text and seed number so you can reproduce or refine the generation.
Tips for Consistently Better Results
Build a prompt library. When you write a prompt that produces excellent results, save it. Over time, you build a collection of tested prompt structures for different types of images, portraits, landscapes, products, illustrations, that you can adapt rather than starting from scratch each time. Many experienced users maintain prompt templates with placeholders for subject, style, and mood that consistently produce strong results.
Study other people's prompts. Midjourney's community gallery shows the prompts used for each image. Civitai posts include generation parameters. Reddit communities like r/midjourney, r/StableDiffusion, and r/dalle share prompts alongside results. Reverse-engineering prompts from images you admire is one of the fastest ways to improve your own prompting skills.
Understand your tool's biases. Every generator has default tendencies. Midjourney defaults to dramatic lighting and rich colors. DALL-E defaults to cleaner, more neutral compositions. Flux defaults to high contrast and sharp detail. Learning what your tool adds automatically helps you write prompts that complement rather than fight against these tendencies.
Generate more than you think you need. AI image generation is cheap (often free), and the quality varies between generations of the same prompt. Generating 8 to 16 images from a well-crafted prompt and selecting the best one produces consistently better results than trying to perfect a single generation through extensive refinement.
Start with Leonardo AI's free tier and a specific prompt describing subject, style, and lighting. Generate multiple variations, refine your prompt based on results, and expect that iteration, not a single perfect generation, is how you produce great AI images.