Which tool to choose

It depends on how faithful you want the new image to stay to the original.

  • For changes guided in words ("turn this room into a modern style"), the generator built into ChatGPT or Gemini accepts the photo and the instructions in plain language.
  • For strong artistic reinterpretations, Midjourney uses the image as a reference for style and composition, with very creative results.
  • To keep the structure almost intact while changing only some elements, you need tools with control over the "strength" of the reference (how far it moves from the original).

How to do it

Image-to-image has two ingredients: the starting photo and the instruction on what to change. The result depends on how far you ask the AI to move from the original.

  1. Upload the reference photo. Look for the upload icon (a paperclip or an image) in the text box, or drag the file. From a computer or from a phone the gesture changes (you drag or pick from the gallery), but the result is the same.

  2. Explain what you want. Be clear about what to keep and what to change: "keep the composition, change the style to watercolor" is different from "use only the colors of this photo".

  3. Adjust how faithful to stay. Where it exists, there's a control over the intensity of the reference: high to stay close to the original, low to give the AI freedom. Without that control, say it in words ("stay very faithful to the photo").

  4. Iterate. Compare the result with the original and correct in words what was lost or changed too much.

If you can't find how to upload the photo, write in the box: "I want to start from a photo I'll upload to you, tell me how to attach it" and follow the instruction.

The working syntax for a clear image-to-image instruction:

I'm uploading a photo. Generate a new image like this:
- keep: the composition and the main subject;
- change: the style, making it an oil painting with warm tones;
- don't add elements that aren't in the original photo.
Stay faithful to the arrangement of the objects.

A concrete example

Chiara has a photo of her living room and wants to see how it would look in a more modern style before renovating it. She uploads the photo into the generator, writes "keep the arrangement of the furniture and the perspective, change the style of the furnishings to modern minimalist, light walls". The AI returns the living room, recognizable but reinterpreted. Chiara asks "put a gray sofa in place of the current one" and gets a targeted variant. She uses it as inspiration, not as a definitive project.

When it does NOT work (and how to fix it)

If the result moves too far from the photo

The AI had too much freedom. Fix: raise the intensity of the reference if the control is there, or write explicitly what to keep ("preserve the composition and the proportions") and what not to touch.

If the starting photo is low quality

A grainy photo limits the result. Fix: start from the best version you have; if needed, improve the photo first with a cleanup tool and then use it as a reference.

If the AI changes faces or details you wanted identical

Keeping a specific face faithful is hard in generic image-to-image. Fix: for portraits use the dedicated tools that preserve facial features, or isolate the change to the background only and leave the subject unchanged.

A tip from someone who actually uses it

Decide beforehand what is sacred in the photo and write it as a constraint. Most disappointments come from not telling the AI what NOT to change: the system, left free, reinvents everything. An instruction that says "keep this, change only that" gives you the control the generator on its own doesn't.

Frequently asked questions

Can I use a photo found online as a reference?

Technically yes, but watch the rights: someone else's photo is covered by copyright, and generating a variant doesn't automatically give you the right to use it. For public or commercial use, start from your own photos or from images with a free license.

Does image-to-image modify my original file?

No. The generator creates a new file and leaves the original intact on your device. You're producing a reinterpreted copy, not overwriting the starting photo.

Does starting from a photo always give a better result than starting from text?

It's the belief to scale back. The photo helps when you want to preserve a precise composition or style. But if you have something new in mind, a well-made text description gives more freedom and often a cleaner result. The photo is a useful constraint, not a universal shortcut: it's needed when the original matters, not always.