All writing

Essay

Photography made me better at directing AI.

I learned photography by shooting, not in a classroom. What that left me with was a vocabulary for images, and a prompt is mostly vocabulary. The same words that direct a crew direct a model.

David Galvis 2026 7 minute read
Bogotá at night, seen from the eastern hills

I never studied photography. I bought a camera, took bad pictures, looked at why they were bad, and took slightly less bad ones. Then people started paying for them, which is a fast way to learn what a technical decision costs. I built a studio in Colombia while I was still at university and shot for more than forty brands before I moved to the commercial side of the table. I stopped shooting for companies. I never stopped shooting.

For a long time I filed that under a former life. It turned out to be the most transferable skill I have, and the transfer showed up somewhere I did not expect: in how well I can direct a machine.

Learning it the empirical way

The empirical route is slower and it leaves you with something specific. You do not learn the rules as rules. You learn them as consequences. Open the aperture and this happens to the background. Move the light thirty centimeters and the face changes shape. Shoot at eleven in the morning and the color goes hard and blue and there is no fixing it later.

Because every one of those came with a picture attached, the vocabulary is not abstract. I can hear a word and see the result. That is the whole asset.

What the camera actually taught

Strip away the gear and photography is a small set of controllable variables. This is what anyone who shoots for a few years ends up holding.

Optics. Focal length and what it does to a face, aperture and depth of field, distance and compression. Why a portrait at 24mm looks wrong and nobody can say why.

Light. Key, fill and rim. Hard against soft. Direction, quality, ratio. Where the shadow falls and what that shadow says about the subject.

Color. Temperature, tint, saturation against vibrance, and how a single grade across mixed sources is what makes a set of images look like one set.

Composition. Thirds, leading lines, negative space, symmetry, where the eye lands first and where it goes second.

Geometry. Camera height, angle, horizon, the lines a building makes and whether they are converging on purpose.

Time. Shutter and motion, the hour of the day, what light does between five and seven.

Six categories. Nothing exotic. But holding them means I can look at any image and say what specifically is wrong with it, in words, in one sentence.

A vague prompt is a vocabulary problem

Which is exactly what a generative model needs. Most bad prompts I see are not bad thinking. They are a person who knows what they want and does not have the words for it, so they write "make it look professional" and get back somebody else's idea of professional.

Without the words

make this product photo look more premium and modern

With them

single soft key at 45 degrees camera left, large source, fill at about a 4:1 ratio, neutral 5200K, shallow depth of field, product at lower third, clean negative space above for the headline

The second one is not a better prompt because it is longer. It is better because every term in it is a variable the model was trained on, so each one moves the output in a direction I chose. And when the result is off, I can name which variable to change instead of rerolling and hoping.

The skill is not prompting. The skill is having the words for what you can already see.

The same words direct people

The unexpected part is that this arrived in my job long before AI did. At Digitech, working with the design team on the material the sales team carries into meetings, the difference between a fast review cycle and a slow one is entirely whether I can say what is wrong.

"I do not love it" costs the designer three days. "The eye lands on the logo before the headline, and the crop is fighting the horizon" costs them twenty minutes. Same feedback, one of them is actionable.

That was true when I directed shoots, it was true running content teams in Bogotá, and it is true now with a model in the loop. Directing is a language problem, and the camera is where I learned the language.

Knowing the programs

The other half is knowing the tools well enough to know what operations exist. Years in Photoshop, Lightroom, Premiere and After Effects leave you with a mental index of what can be done to an image or a cut: masking, frequency separation, curves, keying, tracking, stabilization, transitions, and what each one costs.

That index is what lets me ask a model for the right operation. I am not asking it to make something pretty. I am asking it for a specific move, the same move I would have made by hand, which means I can also tell whether it made it.

Video, and building the tool

Video is where this compounds fastest right now, because the current generation of tools wants exactly this kind of direction. Guiding something like Higgsfield through a shot list is a cinematography conversation: lens, movement, duration, continuity between cuts. Programmatic video with a framework like Remotion is the same thinking in code, where a composition becomes a component and a timeline becomes data.

And past a certain point you stop guiding other people's tools and build your own. An editor shaped around the three things you actually do, assembled with Claude in an evening. Written down that sounds like very little. Anyone who has tried it knows it is far more real than it sounds, and it gets more real every few months.

The part you cannot fake

I mix AI into images the way Christopher Nolan mixes camera work with computer graphics. The machine takes what it does best. Composition, light and structure stay craft, because someone has to decide what the frame is for.

The part that does not transfer to a model is having been in the room. Having stood in that light, at that hour, and known that this was the moment. Everything downstream of that can be generated. That cannot, and I think it is about to become the most valuable half of the job.

David Galvis, seated on a bench
Ten thousand frames taught me the words. The words are what let me direct anything else, including a machine.
David Galvis

Where to start

Three ways to build the vocabulary you are missing.

Name what is wrong before you fix it.

Next time an image is off, write one sentence about why, using a specific variable. Light direction, crop, color temperature, focal length. If you cannot finish the sentence, that is the gap.

Learn one variable at a time, by consequence.

Shoot the same subject at three focal lengths. Move one light and nothing else. The point is to attach a word to a result you saw yourself, because that is the version of the word you can use later.

Give the model a shot list, not a mood.

Frame, lens, light, movement, duration. Treat it like a crew that will execute exactly what you said. That constraint is what forces the vocabulary to get precise.

Why this matters to me

I spent years thinking the camera was a detour from a commercial career. It was the training for it. Photography taught me to see the thing, marketing taught me to explain it, sales taught me to be in the room when somebody tests it, and all three of those turned out to be one skill wearing different clothes.

Back to home

Keep reading