The Constructed Image: Notes on Teaching AI Photo Tableau at Gengdan Institute

The Constructed Image: Notes on Teaching AI Photo Tableau at Gengdan Institute

By Bronisław Kózka

Article content
Image by Gengdan lecturer Bing Zhai (Brian)

Introduction

I have been working between Melbourne and Shanghai for more than fifteen years. In that time my engagement with China has moved through several forms: as a curator at the Pingyao International Photography Festival from 2010 onward, as an exhibiting artist at West Bund Art and Design and Parts Unknown Art in Suzhou, as a residency artist at the Swatch Art Peace Hotel in Shanghai, and most recently as a visiting scholar at the Gengdan Institute of Beijing University of Technology. The work I make and the conversations I have been having about it are now genuinely cross-cultural. The studio in Collingwood and the studio in Shanghai are not separate practices.

My teaching has followed a similar arc. At RMIT I have spent close to two decades teaching photography across the undergraduate programme, with a particular focus on tableau photography, studio lighting, and the kind of constructed image-making that begins from a brief or a fiction rather than from observation. The early work was about what a tableau photograph could be: the production of a fictional moment, often a moment that never happened, photographed as though it had. Over time the teaching has moved closer to the studio practice. In recent years it has become explicitly character-driven, drawing on collaborations with writers and on the methodology I developed for the long-form project Remembering What Never Happened. That project taught me something I now teach others. Once you know who a person is, every production decision becomes specific rather than arbitrary.

The arrival of generative AI has changed what comes next, but it has not changed the underlying logic. If anything, it has made the logic more visible. A generic intention produces a generic image, regardless of medium. A specific intention produces a specific image. AI exposes this faster and more brutally than any other tool I have worked with. The course described in this essay is the first time I have brought all of these threads together for a single cohort: tableau, character, lighting, photographic literacy, and AI as a medium for making images.

I was invited to Gengdan Institute as a visiting scholar with a clear and honest brief. The host asked for a three-week intensive in my area of expertise that would complement the students' existing curriculum rather than duplicate anything they could already access. The course needed to give the students the experience of working with an international lecturer in a contemporary practice they would not otherwise encounter. What follows is an account of what I taught, why I taught it that way, what the students made, and what the course taught me in return.

Article content
Sunshine House, 2008 B.Kozka

The Cohort and the Constraint

The students at Gengdan were not photography students. They were design students, working in an audiovisual unit that was open enough in form to accommodate photography as one of several possible directions. This matters for everything that follows. The cohort came in with strong visual literacy in design conventions: typography, layout, branding, screen-based image, social media. They did not come in with the assumptions that photography students often bring, which can be useful and limiting in equal measure. They had no equipment fetish.

They had no inherited orthodoxy about what counted as a photograph.

Equipment access was limited. Institutional DSLRs were not available at

scale, and asking the students to acquire personal SLR kit for a three-week course was neither realistic nor useful. The deliberate decision, made before the course began, was to treat the phone as the serious primary instrument of the course rather than as a fallback. This is not a compromise. It is a curriculum decision with arguments behind it, and those arguments will appear throughout the rest of this essay. Most of the students globally who will encounter this material in the coming years will be making images with phones and with AI. Pretending otherwise in the classroom would be a teaching failure.

The cohort was Mandarin-first. My teaching is in English. The course was therefore built from the outset as a multilingual environment, supported by translation infrastructure I will describe in the next section. This was not an obstacle to be worked around. It was a condition the course had to be designed for.

The Language Infrastructure

Teaching philosophically demanding material in a language most of the students are not fluent in is a real challenge. It is not a translation problem with a technical solution. It is a teaching problem that translation tools either help with or get in the way of. The course was built on the working assumption that AI-mediated translation, used carefully, has crossed a threshold where serious cross-cultural teaching is now possible without sacrificing depth. This assumption was tested every day for three weeks.

In the live classroom the setup was as follows. A microphone running through the computer fed my speaking voice into Doubao, the Chinese AI tool developed by ByteDance, which produced real-time speech-to-text translation from English into Mandarin. A secondary speaker carried the audio cleanly so the students could hear without competing with ambient room noise. The system had occasional hiccups. Idiomatic English, technical photographic vocabulary, and proper nouns sometimes confused it. None of these failures were structural. At the level of ideas, the translation quality was strong enough to teach genuinely difficult conceptual material without dilution. Students were following the philosophical argument of the course, not just the practical instructions.

Asynchronously, two further pieces of infrastructure carried the rest of the load. The full course schedule, class plans, and reading material lived on my website, kozka.com, under Teaching 2026. The site has a language toggle that renders the entire site, including each individual class page, in Mandarin on demand. Students arrived at every class already familiar with the day's material in their own language. The contact hours were not consumed delivering content. They were spent applying it.

Article content
Mini APP through WeChat allowed real-time upload and feedback

The second piece of infrastructure was WeChat. Through a mini-program file-sharing function, students uploaded works in progress directly into the class group: monologue drafts, prompt iterations, image tests, mood boards. I could read, translate, comment, and paste translated feedback back to the student in something close to real time. The usual asynchronous lag where feedback arrives a week after the work was made effectively collapsed. The student saw my response while the work was still active.

Two observations are worth making on this. The first is that the quality of AI translation at the academic level has crossed the threshold I described above. Before the course began I tested several tools by translating an essay I had written about my own art practice into Chinese and sending the translations to academic colleagues for assessment. ChatGPT, when prompted specifically for an academic register, produced the strongest results. The limiting factor in cross-cultural academic communication is no longer the AI. It is the quality of the prompt and the calibration of register.

The second observation is more specific to this course. The course was about directing AI as a tool for image-making. The teaching of the course was itself an example of directing AI as a tool for cross-cultural communication. The students experienced the method in the form of the class itself, not just in the content of the class. Compared to a course I taught at the same institution the previous year, on different subject matter and with a less developed language infrastructure, the difference was visible in the students' work. Comprehension was stronger. Engagement with feedback was more substantive. The quality of final outputs, judged against what was attempted, was higher.

The Pedagogical Philosophy

Three claims hold the whole course together.

The first is that looking is the first technical skill. Before composition rules, before camera settings, before lens choice, before any prompt is written, the student has to slow down enough to read a frame. Most beginners think photography is about what you point the camera at. It is actually about what you decide to include in the frame, and what you decide to leave out. This is true of documentary, true of tableau, and emphatically true of AI work. The student who cannot read an image will not be able to make one.

The second claim is that pre-production is thinking, not administration. The decisions you make before the shutter is pressed, or before the prompt is typed, are the decisions that determine the image. Every decision you make on shoot day or at the keyboard that you could have made beforehand is time and attention you are burning. This shifts the centre of gravity of an image-making practice. The shoot or the generation becomes the last step in a sequence rather than the first, and the quality of the result follows from the quality of the thinking that preceded it.

The third claim is that specificity is the work. Generic intentions produce generic images, regardless of medium. AI exposes this faster than anything else I have taught with. A vague prompt produces a vague image. A specific prompt, written by someone who genuinely knows their character, their lens, and their light, produces a specific image. The discipline of specificity is the discipline that transfers across every part of the course, from the 30 Questions for character development to the JSON-structured prompt at the end.

These three claims invert a common assumption about what an AI photography course should be. The assumption is that such a course is technical, that it is fundamentally about prompting, and that the prompting is the skill. The course I taught argues the opposite.

Prompting is the last skill in the sequence. Everything before it is what determines whether the prompt has anything to say.

The Phone as a Serious Image-Making Device

Article content
Guide for students using their iPhone in a different way...

The first practical session of the course was about the phone. Not as a fallback, but as the deliberate primary instrument of the work. The mindset shift required is one of the most important things I taught. A phone, for most people, is at least three devices. It is a snapping device for personal use: selfies, photographs of meals, photographs of friends. It is a visual note-taking device: screenshots, memory aids, references collected on the move. And it is, potentially, a serious image-making instrument used with attention, intent, and craft. Most students arrive using it only as the first two. The course asks them to move into the third register, and that move is mental before it is technical.

I produced a bilingual guide for the course, Seeing Through the Lens, which became the textbook for this part of the work. The guide opens with a claim that was unsettling to some students and useful to all of them: Your phone is not a neutral recording device. Every image it produces is shaped by hardware choices, trained AI models, and software decisions made by engineers and data scientists before the phone ever reached you.

This is a critical idea for the rest of the course. AI is not arriving in the final week. AI is already inside Class One, in every student's hand. Smart HDR, Night Mode, Portrait Mode, AI scene recognition, automatic colour grading: every shutter press is already a negotiation with a model. Naming this early reframes the rest of the course. The students are not learning whether to use AI. They are learning to direct AI, beginning with the AI already in their phone.

The guide covered four substantive areas. The first was the phone camera system itself, treated as a multi-lens optical instrument rather than as a single zooming camera. Modern phones contain three or four physical lenses with distinct fields of view: an ultra-wide, a main, and one or more telephotos. Knowing your own phone, knowing which lens engages at which zoom multiplier, is a literacy skill.

The second area was field of view and focal length. This material came back later in the course, when the students began constructing JSON prompts and needed to specify the lens used for an imaginary shot. The principle the guide emphasises, move, don't just zoom, is one of the most transferable lessons in the course. Changing your physical position relative to your subject produces a more interesting image than changing the zoom multiplier. This is composition discipline disguised as a technical tip.

The third area was computational photography. A Night Mode image is a composite of dozens of frames averaged and sharpened by software. It shows a version of the scene that no single moment actually looked like. This is the same argument that tableau photography makes in another form. The image was always a construction. The phone simply makes the construction visible.

The fourth area was depth of field. Phone sensors are physically small and their lenses are short, which produces deep depth of field across the whole frame regardless of the f-number on the spec sheet. Portrait Mode simulates shallow depth computationally and is reasonably honest about its failure modes. Edge detection breaks at hair, at transparent objects, at complex backgrounds. The guide teaches students to read those failures rather than be defeated by them.

The Course Architecture

The published course on the website shows six classes. The course as actually taught had ten sessions, including a final presentation. The three additional sessions were not added at the end. They were interleaved at the points where the conceptual material met its practical limit.

The studio lighting class fell after the first shoot, when the students had a tactile sense of why they needed it. The monologue writing workshop fell after Class Five, when it became visible that the writing required more time than a single session allowed. The shooting and prompting exercise fell between the conceptual introduction of AI and the final project, so the students approached the brief already knowing how a captured image and a generated image could be made to work together. Where each class sat in the sequence was itself a pedagogical decision.

The full sequence ran as follows.

Class One: Learning to Look. The starting point was portraiture, treated not as a soft entry point but as a discipline of attention. We studied a small number of major portraits: Karsh's Churchill, Arbus's child with the toy hand grenade, a careful selection of others. The exercise was recreation rather than imitation. Each group chose two reference images, scouted a location on campus, planned the time of day for natural light, and prepared to recreate the chosen frames. The question driving the work was simple. What is in this frame, and why?

Class Two: Pre-Production as Thinking. Two case studies, drawn from my own practice. A commercial advertising campaign for CareSuper, with its brief, shot list, location scout, and produced result. And a long-form personal project, Remembering What Never Happened, where the same rigour appeared underneath very different conditions. The students went out to scout their chosen locations and produced test frames. The point was that pre-production is not administration. It is the work.

Class Three: The Shoot. The students returned to their planned locations and produced their recreations on phones. The discipline I emphasised was simple and physical. Hold the reference image up next to the camera before every frame, not after. Memory drifts. The reference does not. Three roles ran in each group: photographer, director, subject. We critiqued the results back inside, focused on the compositional question rather than on aesthetic preference.

Class 3A: Studio Lighting Practical. This was the first interleaved class, added directly after the recreation shoot. The students had just produced images in available light and could now see, in their own work, where that light was working for them and where it was not. The studio at Gengdan was equipped with continuous and strobe sources, modifiers, sandbags, and stands. We worked through hard light and soft light, direction (key, fill, back, edge), quality, contrast, and colour temperature. The students recorded the results on their phones. The point was not to send them home as studio lighting technicians. The point was vocabulary. You cannot prompt for cinematic side light from a single hard source if you cannot see, name, or build that light.

Class Four: Tableau as Field, with the 30 Questions. A survey of four practitioners working in different positions on the constructed image: Gregory Crewdson, Jeff Wall, Cindy Sherman, and Shane Hulbert. Then the introduction of the 30 Questions, a structured tool for character development I have used in my own practice for years. The students began building characters. The questions covered identity, world, personality, relationships, fears, desires, regrets, and the physical world the character inhabits. The point of the exercise is not to answer every question perfectly. It is to know the character well enough that every subsequent decision becomes specific rather than arbitrary.

Class Five: Voice, Object, World. The character developed in Class Four needed to speak. We looked at internal monologue in cinema across six examples: Blade Runner, Taxi Driver, The Shawshank Redemption, Chungking Express, In the Mood for Love, and Red Sorghum. The function of the monologue is to give us direct access to the interior life that a face and a body can only hint at. For tableau and AI work, the monologue is scaffolding that the audience never sees. The actor or the AI subject who knows the monologue carries that knowledge in the face, the hands, the eyes. The camera reads it. The students wrote the opening of their own monologue: 150 to 200 words, first person, one specific moment. Alongside the monologue, props, costume, and the physical world the character inhabits.

Class 5A: Monologue Writing Workshop. This was the second interleaved class, and it was added in direct response to need. The first monologue drafts revealed that the writing was the most demanding part of the course for the cohort. Writing in a literary register, often in a second or third language, while constructing a fictional interior voice, is hard work. It would be hard for native English-speaking students at RMIT too. I produced a bilingual Monologue Writing Guide / that distilled the principles down to ten: compression rather than repetition, choose a specific moment, write from inside the mind, use contradiction, let information leak, create movement, write for performance, be specific, leave it unresolved, and work through a practical workflow of review, draft, and refine. The workshop ran through these principles with examples and one-on-one feedback. The improvements between the first and second monologue drafts were the single most visible piece of progress in the course.

Article content

Class 5B: Shooting and Prompting Together. The third interleaved class, and pedagogically the most important. The students went out to shoot, on their phones, in pairs and small groups, with a specific brief. Each captured image was to be the seed for an AI-extended scene. The captured photograph might supply the location, with AI used to integrate a character into it. It might supply the subject, with AI used to construct a fictional environment around them. Or it might be a documentary frame that the student then reconfigured into a tableau through prompting. The point was to dissolve the false binary between photography and AI image-making. Most contemporary practice sits on a hybrid axis, and the students needed to see that on the first day they tried it, not only at the final project.

Class Six: Expanding the Field of AI Photography. A survey of six contemporary artists working with AI as a medium: Shane Hulbert, Boris Eldagsen, Nouf Aljowaysir, Ben Millar Cole, Jake Elwes, and Philip Toledano. Six positions, none of them identical. The students were asked to locate their own work in relation to this map before they began the final project. Then the project brief itself: a monologue video, five still images, and a production book, presented together as one coherent body of work at the final class.

Final Presentation. The students presented their projects in the studio. Five framed prints on the wall. The monologue video on screen. The production book on a plinth in front of the prints. The work was assessed as a complete body of work, with the production book as an equal participant alongside the visual outputs.

JSON Prompts as Production Document

The shooting and prompting class introduced the framework that would carry through to the final project. The principle is straightforward enough to state and demanding enough to apply. A structured prompt is a production document. It demands the same specificity as a director's brief or a 30 Questions document. The opening line of the prompt guide I wrote for the course is the line I would put on the wall of the studio if I could:

Think like a photographer, not a prompt writer.

The closing principles of the same guide are the philosophical spine of the whole course in miniature:

If you cannot describe it clearly, you do not understand it yet. Every word in the prompt should do work. Do not decorate. Construct. Think like a director, not a generator. The image is built through intentional decisions, not effects.

The framework runs through twelve fields. The core idea. Character definition. Character interaction. Spatial construction. Environment and props. Aesthetic direction. Lighting. Colour palette and grading. Lens and field of view. Composition. Surface and texture treatment. Final output style. Each field draws on something the students had already learned in earlier classes. Character definition draws on the 30 Questions. Character interaction draws on the tableau practitioners.

Spatial construction draws on the set construction case studies of Grief and Room 101. Lighting draws on the studio practical. Lens and field of view draws on the Seeing Through the Lens guide. Colour palette and grading draws on the aesthetic exercise that came next.

The move from a natural-language structured prompt to a JSON-form prompt was a small but useful step. The JSON form makes the structure visible at a glance. Every field must be filled. Empty fields are visible.

Vagueness is no longer hidden inside a paragraph. A skeleton of the structure looked like this:

Article content

The 30 Questions and the JSON prompt are the same tool in different forms. Both refuse the generic. Both are production documents. The student who has done the work of the 30 Questions has, in effect, already filled in half the fields of the prompt before they sit down to type it.

The Style and Aesthetic Exercise

Once the framework was in place, the students needed to see it applied. I demonstrated the method with two contrasting examples chosen specifically to show that it holds across very different sensibilities.

The first demonstration was Wong Kar-wai. I decomposed the look of his work into named, enterable decisions. Wide-angle lenses (often around 25mm or 28mm) used in close, characters pressed into the frame and into each other. Practical light sources, neon, tungsten reflections. Step-printing and slow shutter producing trailing motion blur (which translates in a still image to particular kinds of suggested movement and gesture). A palette of saturated greens, reds, and ambers; dense, humid, and cinematic. Atmosphere of rain, smoke, and claustrophobic interiors. A tonal register of time, longing, and missed connections. I built a JSON prompt live in front of the class, filling each field with a specific Wong Kar-wai answer, and we generated the result together.

The second demonstration was Wes Anderson. A deliberately opposite register. Longer focal lengths flattening planes, frontal staging, perpendicular framing. Symmetrical composition with the subject centred, frames within frames. A pastel palette, art-directed, reduced in range with strong accents. The tonal register of melancholy hidden inside formal control. The same JSON structure, an entirely different image

Article content

Side by side, the two demonstrations made the pedagogical argument the course had been building toward. A photographer's or filmmaker's look is not mystical. It is a finite set of named decisions. Once visible, those decisions can be specified in a prompt, in a shot list, or in a production plan. The same is true whether the final image is captured or generated. This is also the moment in the course where the design students' existing strengths surfaced and were validated. They were already trained to read style as a system of decisions rather than as a feeling. The course met them where their existing skills were and connected those skills to a photographic and AI vocabulary.

The students then chose their own reference: a photographer, a filmmaker, an aesthetic, a movement. They built JSON prompts and tested them. The exercise produced some of the strongest individual work of the course.

Article content

How the Components Reinforce Each Other

By this point in the course, every earlier class has become a field in the final prompt or a column on the call sheet.

The Seeing Through the Lens guide gives the spatial and computational vocabulary. The 30 Questions gives the character scaffolding. The Monologue Writing Guide gives the interior voice and trains a kind of writing AI does not produce by default. The studio lighting class gives the lighting vocabulary. The Wong Kar-wai and Wes Anderson demonstrations give the method for decomposing aesthetic into decisions. The shooting and prompting class gives the hybrid practice that connects the captured image to the generated one. The Image Prompts Guide and the JSON exercise give the field-by-field production document that synthesises everything above.

Nothing taught is decorative. The student finishes the course as the director of the image, not as the typist of a prompt.

Article content
Article content

The final presentation took place at the end of the third week. Each group produced five framed still images, a monologue video, and a production book. The work was hung on the studio walls in groups, with the production books displayed alongside.

The range of subject matter and genre across the cohort was wide. One group produced a contemporary ensemble set in a bookshop, with characters interacting across reading rooms, library shelves, and a doorway scene that recurred across several frames. Another produced a courtroom narrative crossed with what appeared to be a historical or museum thread, the same lead character moving between a wood-panelled tribunal, a stone-walled cell, and a glass-cased institutional space.

A third group produced a period detective narrative with the title The Red Migratory Birds: A Cross-Time Pursuit of Truth. The lead image was effectively a film poster, with four supporting frames extending the story: a wetland encounter, a cobblestone street with figures gathered around a lamp, a group around a study table, and a tableau of figures bent over papers and instruments. The mise-en-scène, costume, and lighting were coherent across all frames. This was a fully realised genre world.

Article content

A fourth group produced a series titled [IMAGE:, the Red Migratory Birds panel.] , Tradition and Trend Coexist, anchored on a single woman in a green qipao engaged in embroidery. She appeared in four locations across the series: a lantern-lit interior, a balcony at sunset with a guitar-playing partner, a rain-soaked courtyard, and a domestic interior with ceramic vessels on a low table. The same character held across very different lighting conditions and emotional registers. This was the clearest evidence in the cohort that the 30 Questions had produced a person who could be photographed in different conditions and remain recognisably the same person.

A fifth group worked in a contemporary social register: a bakery, a kitchen, a gallery, a wedding banquet, and a moment of conversation between a flower-bringing visitor and a seated figure. A sixth produced a film-set self-reference, with characters meeting in cafes, conducting interviews, and standing in front of clapper boards and tungsten lights. The Chinese banner across one image read, a wrap-party congratulation for a film called Ferry Crossing Bookshop. A group of design students making AI images of a film crew making a film: the course's argument about constructed image-making turned reflexively back on itself.

Article content

A seventh group worked in a more interior, magical-realist register. The recurring object was a blue scarf. The lead character sat in bed reading with a cat, encountered a horned figure performing what appeared to be a ritual in a candle-lit library, and appeared in close-up portraits and over-shoulder shots that worked at the level of mood rather than narrative. This project sat closest to the Remembering What Never Happened methodology I had taught: one character, one recurring object, atmosphere as the primary carrier of meaning.

Article content

Several observations emerge from looking across the cohort.

The character-first method produced legible worlds. Each group had clearly settled on a character or set of characters who recurred across their five frames. The brief produced bodies of work, not collections of single images. The students understood that an image series is a world, not a portfolio.

Production design and styling held within groups. Wardrobe was consistent. Recurring objects appeared. Lighting choices supported character rather than fighting it. The discipline taught through the prompts guide and the studio lighting class transferred into practice in a visible way.

The poster and title-card instinct that surfaced in some of the groups is worth naming honestly. Two of the projects produced what are essentially film posters or promotional layouts. This is design students reaching for the format they know best. Their existing visual literacy as designers shaped the form their final work took. That is a feature of teaching this cohort, not a bug.

Cultural specificity surfaced naturally. The qipao, the Chinese signage, the genre conventions of the period drama, the bakery setting, the Chinese-language poster work. None of these were prompted by the teaching. They emerged because the students were given a method that does not prescribe content and were free to bring their own world to it. This is, I think, the strongest possible evidence for the argument that specificity is local. Teach the method honestly and the cultural particularity of the maker becomes the substance of the work.

Where the Tools Held and Where They Did Not

There is one finding from the course that needs to be stated as plainly as possible. The monologue writing was the strongest part of the course. The monologue video was the weakest. Both statements are true at once, and the relationship between them is genuinely interesting.

The writing succeeded for clear reasons. The combination of the 30 Questions, the Monologue Writing Guide, the dedicated workshop in Class 5A, and the WeChat upload-and-feedback loop produced character voices with real specificity and emotional weight. Many of the resulting monologues were the most articulated work in the course Some were excellent. The students refined them through several iterations, and the iterations were visibly responding to translated feedback that arrived while the work was still hot.

The video struggled for reasons that were partly outside the students' control. Two practical constraints, named honestly. The consumer-tier subscriptions available to the students capped individual AI video clips at five seconds. A monologue video of one to two minutes had to be assembled from a sequence of short clips. Continuity, character consistency, and lip-sync drift between clips compounded across the run. The second constraint was time. Three weeks is short for what the course attempted. The video deliverable arrived at the point in the course when the students' AI tooling literacy was still developing.

The students who succeeded with the video component did so by routing around the constraint rather than fighting it. Some used short clips of the speaking face combined with voice-over against other imagery. Others used a back-of-head or partial-face approach, allowing the audio to carry the meaning. This is a real cinematic strategy with a long history. Wong Kar-wai uses it. Terrence Malick uses it. Chris Marker built La Jetée almost entirely from stills with a single moving moment. The students who arrived at this solution were working in a tradition, whether they knew it or not. The constraint produced a more sophisticated result than the apparent path of least resistance would have done.

For the next iteration of the course, this finding becomes a brief rather than a workaround. I would now ask students to plan for approximately fifteen seconds of lip-synced face, with the remainder of the monologue carried as voice-over against other imagery. This is not a compromise, It is a better brief, It teaches students to think cinematographically rather than to fight the rendering tool.

What the cohort achieved in three weeks, working with consumer-tier tools at the edge of current AI video capability, was substantial. A course twice as long would have given the video work more time to mature. The course was built for the time available, and within that time the students produced strong character work, strong written work, strong still images, and a real engagement with the limits of current tools. That last point is itself a course outcome. The students now know where the edge of consumer-tier AI video lies in 2026 because they have stood on it.

Article content

Reflection: What This Approach Argues

A small number of threads run through everything described above. They are worth naming directly as the reflection that closes this essay.

Slowness is a discipline, not a default. AI rewards speed. The course deliberately rewards depth. The student who knows their character, their lens, their light, and the world the character inhabits produces specific images. The student who does not produces generic ones, regardless of the tool they use. The course's argument is that the specificity has to be built before the image is made, in the writing, the questions, the props lists, the location scouts. The image is the last thing.

Photographic skill has not become obsolete. It has become more valuable. Reading a frame, building light, choosing a lens, directing a subject: these are exactly the skills that allow a person to direct AI rather than be directed by it. The course makes the case that the deepest skills of photography are the ones that survive any change in tooling. Lens vocabulary becomes prompt vocabulary. Lighting literacy becomes prompt literacy. Pre-production discipline becomes the production document that the prompt is.

Article content

The phone, treated seriously, is sufficient. The constraint of the cohort produced clarity. A class that taught DSLR craft in a three-week intensive would not have transferred as cleanly into the AI work as the class that actually ran. Phone plus AI is the realistic toolkit for most contemporary makers globally, and treating that toolkit seriously in the classroom is not a compromise. It is a curriculum decision with arguments behind it.

Article content

Cultural specificity is local, particular, and rooted in the world the maker actually inhabits. Working at Gengdan with Mandarin-first design students drawing on Chinese cinematic and literary references reinforced this. The bilingual provision of the course documents was part of the argument, not an afterthought. The student work demonstrates the argument better than I could state it. The qipao project, the Chinese-language poster work, the bookshop and bakery worlds, the period detective narrative: these emerged from a method that does not prescribe content, applied by makers who brought their own world to it.

Translation infrastructure has crossed a threshold. The classroom can now be a multilingual environment without sacrificing depth, provided the teacher actively directs the tools. Doubao for live speech-to-text. WeChat mini-programs for asynchronous feedback. ChatGPT for academic-register translation. Squarespace's site-wide language toggle for course materials. None of these tools are perfect and all of them have failure modes. Used together, with attention, they made a course possible that would not have been possible five years ago.

Article content

Working at the edge of current tool capability is itself a teaching opportunity. Students who succeeded with the video component did so by directing the medium rather than fighting it. The students who produced excellent five-second clips and used voice-over against other imagery understood, by the end of the course, that directing AI is a craft of knowing what the tool can and cannot do.

The complementarity of an international visiting role lies not in transplanting a home institution's curriculum but in offering something the host's standing curriculum does not. The course worked at Gengdan because it did something the students had not encountered before, taught in a way that brought their own design instincts forward rather than asking them to put those instincts aside. That is the argument I would offer to other artists and educators considering similar invitations. The value is in the complementarity.

The AI image is only as specific as the thinking behind it. The course is about the thinking.ł

Bronislaw Kozka



Bronisław Kózka is an artist working between Melbourne and Shanghai. He is a full-time Lecturer at RMIT University's School of Art and a visiting scholar at the Gengdan Institute of Beijing University of Technology. His practice spans photography, AI-mediated image-making, and the constructed tableau, and has been exhibited internationally at venues including West Bund Art and Design, Parts Unknown Art in Suzhou, the Pingyao International Photography Festival, Arsenale di Venezia, and the National Portrait Gallery London. He is a Hasselblad Master, a board member at Magnet Galleries Melbourne,founder of Down the Lane Studios in Collingwood, and is completing a practice-based PhD at RMIT

Previous
Previous

There is something particularly meaningful about having an article on residencies…

Next
Next

For some time, I have been working on this idea that began as a digital sculpture…