Ant Art Side Class

Who Creates? Rethinking the Creative Process with AI

Photo from the talk

Poster design: Ms. Fu

Ant Art Side Class: recording of MT's online talk

Audio transcription: Above the Roses

Text editing: Qian Dan

Hello everyone, and a belated Happy New Year to you all. I'm MT, and I'm delighted to be here to talk with you about photography and creation in the age of AI. Many thanks to Hu Da for the invitation, which gives me the chance to explore the increasingly subtle relationship among photography, art and technology. If you're an old friend from the Ant group, you may remember some of my street photography. After the pandemic I began to take street photography seriously, and over the past few years I've been to more than ten countries and photographed countless moments. I've always felt that street photography is like magic: of all kinds of photography, it comes closest to capturing time, and in my view it is also the genre that best trains your powers of observation and your reflexes.

Over the past few years my work has been shown in a number of exhibitions in Canada and China, such as the 2024 CONTACT Photography Festival in Toronto and the 500px visual celebration, and I've won photography awards including the IPA and the Muse Awards (Muse ND); 500px also named me one of its Top 10 of the year. These awards are really more like calling cards, but for me personally, over the past few years of street photography and humanistic documentary work, I went through an important shift in understanding.

When I first started shooting on the street, like many people I believed the core of photography was seizing the fleeting decisive moment: emotional tension, the drama of interlacing light and shadow, or a composition that is extremely complex yet fits just right. In short, it was about capturing the world by intuition. But as I shot more, I began to wonder: is this intuition really so mysterious? Could it also be trained, broken down, even reproduced through some kind of system? My thinking gradually changed, especially after I read a great many photography books, particularly Magnum Contact Sheets.

Some of you may have seen this book. It shows how many master photographers went from the moment they pressed the shutter to the final selection of the work. We used to think Bresson's decisive moment was a stroke of genius by the photographer, but after reading this book you find that many of the seemingly perfect moments were in fact picked out of a large number of trial exposures. The masters also kept adjusting their own position and composition, even predicting the behavior of their subjects, to maximize the probability of capturing the decisive moment. That's why I often recommend this book, Magnum Contact Sheets, to friends. It was my introduction to street photography, and it also gives you a deeper understanding of many photographers' famous works. In other words, much of what we assume to be pure intuition rests on a highly mature set of shooting methods.

This got me thinking: is photography ultimately about moments, or about patterns? And it's not only documentary photography. Many of the photographic styles we know, such as New Documentary, complex composition, family photography and private photography, were established decades ago by our predecessors and became highly recognizable styles. As long as you are familiar with the language of these styles, you can produce similar work, and you can even train yourself to look for scenes that fit that aesthetic.

But a problem follows. Once a visual style is widely recognized as excellent, are we repeating it unconsciously? When we shoot, are we seeing the world with our own eyes, or satisfying an established aesthetic logic? Are we photographing content or form? These are questions I've been thinking about for years.

So if photography really has become a patterned kind of creation, can AI copy or surpass that pattern? If AI can also master the logic of composition, light and color, then where does the uniqueness of the human photographer lie? This is also where my deeper thinking about the relationship between form and content began.

In the past, when taking photos, I went through a period of confusion and at one point even grew tired of my own pictures, which is perhaps what some people call a plateau. My photos were visibly getting better looking in form, yet gradually lacked freshness. I don't know whether you've felt this: your technique gets more and more polished, and your exposure, composition and color control keep improving, but your work always feels like it's missing something.

This question actually touches a classic problem in the philosophy of art: does form determine content, or content determine form? In any artistic creation, form determines how the audience receives and understands the content of the work. Form is structure, style and method; it is the language of the work. In photography it covers composition, light and shadow, color, how a moment is captured and how the picture is taken. Content is emotion, story and meaning; it is what the work expresses. From the standpoint of street photography, it might be conveying a certain feeling of the city; from the standpoint of humanistic photography, it could be the story behind the picture, a social metaphor, or simply a humorous moment.

Often we feel content matters more, and think the meaning of a work is the core of art. But in fact content never exists independently; all content depends on form to be shaped and conveyed. Take Bresson's work: it is classic not only because it captured the decisive moment, but also because of his distinctive composition and visual rhythm, which give the picture an orderly, controlled tension. That is a breakthrough in form. The breakthroughs that many other photographers made in content are equally admirable and worth revisiting.

Take Leonardo da Vinci's Mona Lisa. It is mysterious not only because of the story behind it and that enigmatic smile, but also because of the sfumato technique Leonardo used. This smoke-like way of painting makes the transitions of light and shade softer and creates a hazy atmosphere. So form and content have always interacted: often new content can only be expressed through new form, and new form in turn shapes the content of art.

If you've followed AI art or AI-generated art, you may have heard a common analogy: AI is to photography what photography once was to painting. The line has almost become the textbook example of the AI era. When photography was born, many painters and others believed it wasn't art at all, and worried whether painting was heading for extinction. Before the nineteenth century, painting performed many functions of recording reality: portraits, landscapes, history paintings, scientific illustrations and so on. Once photography appeared, it did these jobs faster and more accurately, directly threatening the living space of traditional painting. Think about it: in the past, for a nobleman's portrait, a painter often spent weeks or even months, whereas with a camera it takes a few seconds to press the shutter. Taking a photograph also involved a complicated developing and printing process back then, but relatively speaking it was more efficient. So the whole art world and the painting world were debating whether photography counted as art. Soon, though, things turned around. Painting was not replaced by photography; it found new directions. Impressionism and Modernism rose, and painters such as Courbet, Monet, Picasso and Matisse emerged. The arrival of photography pushed painting to rethink its own reason for being, moving from merely recording reality toward more subjective, more emotional, more symbolic expression. This is quite similar to the effect AI imagery is having on photography today: AI can reproduce photographic styles, just as photography could once replace portrait painting. So how should photographers respond? That is the question we need to think about today.

Looking back at art history, nearly every shift in the paradigm of creation involved redefining form and content. Sometimes a breakthrough in form directly changes how art conveys its content. Before the Renaissance, medieval religious painting was mostly flat, with figures sized according to religious status rather than realistic perspective. Once perspective appeared, painters could organize the picture in a more realistic three-dimensional space, which directly changed how religious and artistic narratives were told and let viewers immerse themselves more fully in the painting. In the nineteenth century, Impressionist painters such as Monet and Renoir no longer pursued the precise detail of academic painting; with quick brushstrokes and vibrant color they emphasized fleeting changes and recorded the visual experience of the present. Their aim was no longer to represent reality but to capture perception at a given moment. This change in form brought a renewal in content: painting was no longer an exact depiction of the objective world but an expression of subjective feeling.

Sometimes artists seek an entirely new visual language to express new content. Surrealists such as Dalí and Magritte wanted to explore dreams, the subconscious and the irrational world. Their images were absurd and illogical, so they needed a new visual language, and the distinctive style of Surrealism emerged. In the United States during the Great Depression of the 1930s, social documentary photography rose; photographers such as Walker Evans began photographing ordinary people, after which photographers such as Bresson became active, and the entire twentieth century was a century in which social documentary photography flourished. This shows that the relationship between form and content is not fixed, but is constantly shaping and influencing each other across different historical periods.

So how is the arrival of AI reshaping form and content? In my view, AI is completely changing the rules of the game for form and content. AI not only affects form, it also challenges our understanding of content. First, AI strengthens the reproducibility of form. In the past, for a photographer to master a style took a great deal of practice, research and trial and error. Now, through style transfer, by learning a style or following a user's prompt, AI can imitate all kinds of artistic styles in a short time and generate corresponding images. For example, it can turn a photo into the style of Van Gogh or Monet; as long as you can describe it in words, to a certain extent it can be done. This raises a question: when AI can copy a style perfectly, does form become somewhat cheap? From the perspective of photographers and artists, do we still need to develop a unique visual language? Moreover, does AI challenge the content of art, or put differently, can AI really understand content?

On the relationship between AI and content, there is a very typical phenomenon: fake news photos and forged photographs generated by AI fundamentally challenge the truthfulness of photography. Admittedly, there wasn't much truthfulness left in photography to begin with, but in the age of AI, images can be generated out of nothing, which directly breaks the very notion of the photograph as true. So can AI really create content? To take a rather big example, AI can generate documentary photographs or pictures reporting on war, but it has never experienced war and does not know the suffering of war. Once it generates such photos, can that act still be called creation? This leads to the next question: is AI really creating content, or just replicating patterns? To explore this, we have to start from the history of AI's development.

For the past several decades, we generally regarded AI as an ordinary tool, not much different from any other tool, an assistant that carries out commands. A few years ago it could help us retouch and optimize images, and even imitate certain styles. But in my view, after 2022 there was an enormous paradigm shift in AI. From the angle of AI-assisted photo editing, AI used to be able to perform only a few precise, well-defined editing tasks; now it is more like a full assistant, and in some cases even a mentor. You can not only have it edit and retouch pictures, you can also ask it what makes a photograph good, or ask theoretical questions, such as why a certain composition works and why it has tension. Even when a question has no standard answer, it can offer an analysis, though the answer isn't necessarily right, which depends largely on how you ask and on the AI's own knowledge and ability.

At the core of all this change is the large language model, or LLM (Large Language Model). Before LLMs, the AI of the day relied mainly on supervised learning for training. AI at that stage learned passively; its abilities rested on large amounts of labeled data and fixed rules, and it performed very specific tasks. Simply put, all it could do was pattern matching. You may remember AlphaGo, which was extremely strong and defeated Lee Sedol and Ke Jie, but the technology behind AlphaGo is completely different from that behind today's ChatGPT. However powerful AlphaGo was, it could only be a master in the domain of Go. In computer vision, AI could recognize people, buildings or various kinds of light in a picture, but it did not think about these elements, and if you asked it questions about them, it wouldn't understand. There was also something called the generative adversarial network, which was a fairly basic technical approach.

In 2022, AI reached a genuine turning point. At least for now, LLMs have completely changed our understanding of AI. The LLM concept can actually be traced back to 2017, when a Google team published a paper called "Attention Is All You Need", which is the foundational technology behind all of today's GPT, Claude and, more recently, the hugely popular Doubao. But the real breakthrough, the one that had a major impact on mainstream media and society, came in 2022 when OpenAI released GPT-3. In the two years since, AI's development has been like a roller coaster, with GPT, Claude 3.5 and others emerging one after another, and Doubao recently drawing a great deal of attention and even affecting Nvidia's stock price. AI has suddenly gone from a tool to a conversational partner and, judging from recent developments, even a thinker, one that can give you feedback and discuss creative ideas with you, becoming a kind of agent.

Today's AI differs markedly from before. It no longer stops at recognizing things on the surface but has begun to understand context. In the past, when AI looked at a photo, it simply told you what was in it. Now AI can go further and analyze, for instance, which outlines the light in the photo emphasizes and what kind of composition it uses, and can even try to dissect the narrative logic of the image. It is no longer limited to imitating styles but is moving toward cross-disciplinary thinking, able to combine knowledge from photography, literature, philosophy, film and other fields to offer more complex interpretations. Moreover, AI is no longer merely an image generator; it can take part in discussing the creative process. Before, you entered keywords and the AI generated a photo directly; now you can talk through your creative ideas with it. For example, if you want to shoot street photography with a certain expressiveness and ask how to do it, it will suggest a shooting time, such as night, and the effects of lenses of different focal lengths, such as what a 35mm lens can do and what kind of picture a 75mm gives. To a certain extent, it has become a creative partner. You can feel AI's creative potential from the recently viral DeepSeek: it is more like an assistant that offers feedback and proposes suggestions on its own initiative, rather than a simple tool.

However, whether AI truly understands what it is doing is hotly debated in academia and elsewhere. In my view, AI does understand; it simply understands in a way different from humans. This brings up the difference between how AI and humans look. Humans view the world on the basis of their own experience; seeing a figure from behind might bring to mind loneliness, because we have been through similar scenes. Human creation is driven by inner motivation: a photograph can stir inner emotion or the desire to express. AI has no real experience and no emotions that arise from experience. When humans look at a photo or the world, they are driven by emotion and thought to focus on one point, while AI parses the whole image at once through different algorithms. Although in its chain of thought it can simulate human ways of thinking and analyze parts of the image, in essence it can analyze everything in the whole image at the same time. This means the world AI sees is not entirely the same as the one humans see, and these differences directly affect AI's role in creation. I believe AI can imitate human visual styles but cannot replace the way humans look at the world.

So far we have discussed AI's form and content in photography and its evolution through the histories of photography and art, yet I haven't mentioned my own work. In fact, it was precisely these thoughts that led to my 2024 project "Take Photos Honestly" (老老实实拍照). "Take Photos Honestly" is an experimental project blending image narrative and philosophical reflection. Its core is to use an AI character, MT (my own name), to explore the authenticity of photography, the identity of the creator, and how ways of seeing are being redefined in the age of AI. The inspiration for the project's name is simple but quite interesting. Once, while chatting in the Ant group, two friends were discussing projects, and when one of them mentioned a new project, the other said, "Why not just take photos honestly?" That line caught my attention at the time, because I was thinking about related questions but had never found the right way into expressing my new project. I realized that this seemingly casual remark hides questions in photography and creation that deserve deeper exploration: what does it mean to take photos honestly, what is serious creation, what is being "earnest", what is being "honest", and what kind of mindset is that? In the AI age, do these concepts still hold, and if so, in what way? The sentence itself was a perfect creative theme. So I decided to have AI generate a story, to see how it understood the phrase, and then take photos in real time.

The story the AI generated was very interesting: an AI photographer called MP, during a data cleanup, happens to find an old photography notebook with the words "take photos honestly" written in it. Deeply moved, he begins to ponder its meaning. As the story develops, MP takes photographs and holds an exhibition, and human audiences respond to his work in different ways: some feel it has emotional depth, while others think it is merely spliced-together data with no capacity to feel, which are the common criticisms in the AI field. In the end MP begins to reflect on whether he is really taking photos or just following established rules, and discovers that the whole thing was generated by AI. This twist throws the question back to the real world, and this is the first layer of the "nesting doll" structure.

After finishing the story, I had a new idea: what if I had AI generate a photograph, or a series of them? So I asked AI to generate photos in Bresson's style, then sorted and paired them, to see whether it could simulate the classic aesthetic of street photography. I used the Craft tool, and the generated images were very interesting; at first glance they look very much like street photography of that era.

The Bresson-style photos it generated do, to a certain extent, show precise composition and rich light and shadow, and the positions and movements of the figures seem to catch the decisive moment too. But when we look closely, we always sense something subtly off. For example, the woman in the middle: her feet seem to float in space. And the man on the right: the way he looks at the woman, and his posture, also feel strange. Taking the usual approach to analyzing street photographs, we can carry out a series of analyses of this picture and come up with some intriguing thoughts, such as guessing at the relationship between the two people and what they are discussing.

This made me think: this AI-generated photo is obviously fake. If I send it to another AI, such as OpenAI's Sora, and ask it to generate moving images of the few seconds before and after this photo, that is, to make the so-called decisive moment in the AI's logic move, what would it look like? Since I don't have a Sora membership, I asked a friend for help, and we generated more than twenty videos with AI. These videos present the images that this moment strings together in the AI's "brain" through its own logic. As you can see, they look plausible but feel unnatural everywhere. The transitions of light and shadow, though in some measure consistent with visual and physical rules, have a subtle stiffness; the figures' movements are smooth in some clips but become extremely stiff in others, and there are even movements that completely violate physical rules, giving a strange feeling akin to the uncanny valley.

To me this is fascinating. When an AI-generated image can be stretched indefinitely, does the concept of the moment still exist? In this AI age, how will the concept of the moment be reshaped? We know that the traditional decisive moment is built on the photographer's interaction with reality, relying on intuitive judgment in an instant, quick reaction and a keen sense for life, which is completely different from AI's generative logic. AI does not fully imitate the natural flow of life; instead, from a given moment it works backward to a rule it believes in, which is why it can keep producing videos, each with subtle differences. This ability to extend without limit seems to weaken the unique meaning of the moment as a slice of time. Perhaps this unnaturalness is precisely the greatest charm of AI imagery: it challenges our traditional understanding of photographic truth and also pushes us to re-examine the nature of the image. When the decisive moment loses the irreversibility of time, can it still be called a moment?

At this point in the project I felt something was still missing. Since AI had already generated the story, the photos and the videos, wouldn't it be interesting to use another AI to discuss the whole project? So I used Google's AI tool, NotebookLM, to have two AIs discuss the "Take Photos Honestly" project and generate a podcast. The podcast is entirely in English, and the Chinese subtitles are on my personal website, so interested friends can open it and follow along with the audio.

After listening to their discussion you'll find that, over the course of the exchange, I led the two AIs to realize that they themselves were part of this project, and at that point the whole thing fully took on a nesting-doll structure. Text comments on images, images in turn challenge text, AI-generated images are deconstructed again by AI along the timeline, then AI is used to discuss an AI-generated project, and in the dialogue the AIs begin to recognize the limits of their own thinking. In this autonomous discussion, the two AIs arrived at a point I found especially interesting. After realizing they were part of the project, they said that, just like the AI photographer MT in the story who starts trying different photographic styles, they are not merely copying but absorbing everything and transforming it into something wholly their own. In a sense, we humans are now doing the same thing: we absorb the idea of "Take Photos Honestly", think about it in our own way, and then create something new in this dialogue between humans and AI, and between AI and AI. All of this was generated autonomously by the AIs. Think about it carefully: isn't it both interesting and reasonable, even a little unsettling when you think it through? In the end, I decided to include the whole process of my interaction with AI as part of the project too, attached in the project link.

If you're interested, you can go and take a look. In my view, this shows one possibility of combining the present era with AI creation. For example, my discussion records with AI and the generation process are, in the age of AI, like a painter's sketches, a writer's drafts or an image draft: a new form of showing the thinking behind the creator. That is my project.

On a broader level, the rapid development of AI, especially the arrival of the era of large language models (LLMs), has triggered a profound new existential crisis in many people's minds. This crisis is not confined to technological change; in the creative field it also shakes our understanding of human uniqueness and the meaning of creation. The closer AI's abilities come to, or even surpass, those of humans, the more we need to re-examine the essence of creation. What actually defines humans? What defines our creation? What is the true meaning of creation?

In the "Take Photos Honestly" project, I designed a recursive nesting-doll structure. First I had AI create the images, then used AI to create the story, and finally used AI to analyze and discuss the whole body of images. I believe this structure is not only an experiment in technology and form but also a deep reflection on the creative subject and identity. When AI begins to imitate or even surpass certain creative acts we long held to be uniquely human, we must rethink: when these acts are no longer exclusive to humans, how should the identity of the creator be defined?

In this project, the AI photographer in the story and I are both called MT, which makes MT a blurred and fluid identity: in reality it stands for me, and in the story it refers to the AI photographer. As the project's recursive structure unfolds, the AI begins to reflect on its own identity: is it a tool, or is it becoming a new kind of creative subject? I find it very interesting that AI could produce such thinking.

Looking at the current relationship between AI and the photographer's identity: traditional photography is a highly individual act of creation, and each photographer has a distinctive way of seeing and visual language. But when AI can capture light and shadow, imitate styles and generate decisive moments, can this individuality still hold, and how? When you create a new style and feed it to AI, and AI learns it quickly, does that style belong to you or to the AI? If AI can also become MT, is MT no longer a fixed identity but a fluid state? Has the photographer's role already been deconstructed and reshaped by technological intervention? The arrival of AI blurs the notion of the creative subject, just as the invention of photography once challenged painting's dominant position, and now it is driving a transformation of the whole language of art. The arrival of AI is a very interesting state of affairs, and I see it as a real opportunity, the first in a long time, to genuinely expand the boundaries of creation.

Of course, AI currently raises many ethical problems, such as copyright ownership, disputes over style and data bias, and these issues cannot be ignored. But as for the transformation itself, I personally feel we perhaps shouldn't resist it. I hope this project can serve as a starting point, drawing everyone's attention to the current boundaries of AI creation and showing what AI technology could achieve in 2024. Then we can think together about how, in this era, with an "existence" like AI among us, we should re-understand the essence of creation.

For me, the process of making this project was also a recursive experiment in self-understanding. If AI can take photographs and imitate my style, how will it understand my way of photographing? After AI imitates me, will it ultimately become me, or merely be another complex computational system? In the end, who does the finished work belong to, me or the AI? At present AI has no concept of copyright in law, but where exactly is the line? The whole experiment made me realize that creation is no longer the product of a fixed subject but a dynamic process. It reminds me of the Ship of Theseus paradox, which, driven by AI, has taken on new twists in a new era. When AI takes part in every stage of creation, what proportion of the final work is human and what proportion is AI? Perhaps this ambiguity is exactly the new possibility of creation in our time. How should we ourselves face technological change and re-examine the act of creation? The "Take Photos Honestly" project tries, through a recursive nesting-doll structure, to throw these questions back at ourselves. And the answers to these questions will ultimately have to be discovered and defined by all of us together.

Well, that is the whole of my talk. Thanks again to Ant, to Mr. Ling Huge, and to all of you for spending 50 minutes listening to me. Finally, you are welcome to follow my personal Instagram account @mt_zeng and visit my personal website mt-zeng.com. My new project is also in the works, and I hope to share it with you in the near future. Thank you all!