Anime characters' eyes are sometimes drawn much larger than real human eyes. In some art styles, they occupy about a third of the face's height. Even so, viewers still recognize the figure as a person and can discern the direction of the character's gaze and even subtle shifts in emotion.
This article takes the expanded expressive range of large eyes as a starting point for considering how works and their audiences change each other. People develop interpretive habits by encountering works. As they discuss, imitate, and create, those habits shape later works, which in turn transform the habits themselves. A work not only communicates through conventions that its audience already shares; it can also test and change those conventions. Focusing on this dynamic, we frame art as "research and development in communication." We then apply this cycle to creation with generative AI and ask where creative work still requires time and judgment once ideas can be turned into drafts with less effort.
Our starting point is a pattern in how people read emotions from faces. Although people draw information from the entire face, changes in certain regions, including the eyes and mouth, provide important cues.
In experiments using photographs of real faces, the most informative facial regions vary by emotion: the area around the eyes tends to matter more for recognizing sadness and fear, while the area around the mouth tends to matter more for happiness and disgust[18]. Where people look also varies with the task, photograph, and observer, and some studies find that the first fixation falls near the center of the face rather than on the eyes themselves[19]. The research therefore supports a limited claim: people do not look only at the eyes, but changes in regions such as the eyes and mouth help them read emotions and mental states. These studies used photographs of real faces and did not test whether the same findings apply to anime faces. Studies showing that infants prefer face-like patterns may indicate only an interest in faces as a whole, so we do not treat them as evidence that artists draw large eyes for this reason.
From this perspective, we can view large eyes as an artistic device that makes subtle changes in expression easier to depict. Because the eyes occupy more space, artists can depict the direction of the gaze, highlights, how far the eyelids are open, and welling tears more distinctly. Although the result departs from the shape of real eyes, it makes visual cues associated with emotion more prominent.
Artists draw large eyes for many reasons, including the traditions of particular art styles, the desire to evoke cuteness or convey a character's age, and production constraints. Rather than looking for a single origin, we focus on how current drawing conventions affect viewers' perceptions and emotions. Based on how artists use the extra space, we infer that large eyes make emotion-related details easier to depict distinctly. This inference is not based on an experiment directly showing that larger eyes improve the accuracy of emotion recognition. In animation, head movement, eyebrows, the mouth, voice, timing, and camera work also convey emotion, so the eyes alone cannot explain everything. Related research has examined how caricatures that exaggerate facial shape or color affect familiarity judgments for familiar and unfamiliar faces[20]. The strongest benefit appeared when participants judged unfamiliar faces whose shapes had been exaggerated. That research did not measure emotion recognition and used faces derived from photographs, so it does not demonstrate the effect of large anime eyes themselves.
Small details such as highlights also stand out more in large eyes. For example, simply removing the highlights without changing the outline of the eyes can make a character suddenly appear drained of life (Figure 1). In real eyes, the appearance of highlights depends on lighting, the orientation of the face, and the viewer's position. In anime, by contrast, artists can remove the highlights without changing any of those physical conditions and use their disappearance as an emotional cue.
Highlights originally represent reflected light. What their disappearance conveys depends on the broader context of the work, the particular scene, the color of the eyes, how the eyelids and mouth are drawn, and the accompanying sound; there is no fixed formula such as "loss of highlights = despair." The point here is not to assign the device a single meaning, but to explain its mechanism: drawing the eyes large makes fine details such as highlights easy to distinguish and usable as emotional cues. This explanation is an interpretation based on how the images are drawn, not on a perception experiment that varied only the presence of highlights.
Figure 1: Large eyes with and without highlights
The image on the left is the original illustration of imos-chan; the version on the right has the highlights removed from both eyes. Because the eyes occupy a large part of an anime character's face, the right-hand version looks noticeably darker.
The example of large eyes shows that departing from realism does not necessarily reduce the amount of information conveyed. Anime omits skin texture while exaggerating the eyes and adding lines and symbols that do not exist in reality. Omission, exaggeration, and addition alter the cues available to viewers. Artists choose what to omit, emphasize, or add according to the work's purpose and intended audience. Scholars have studied manga's lines and symbols as elements of a visual language governed by combinatorial rules[1, 2, 3].
This mechanism is not limited to characters' eyes. The clearest example of adding something that does not exist in reality is the use of speed lines in manga. Just a few lines added to a still image can make a person or object appear to move (Figure 2). How readers interpret speed lines also depends on their experience with manga.
One experiment compared manga panels with speed lines pointing in the conventional direction, panels without lines, and panels with reversed speed lines. Viewing times were shortest for panels with conventional lines, longer for those without lines, and longest for those with reversed lines; EEG responses also differed across the three conditions. The effects of manga-reading experience varied by measure. More experienced readers tended to spend longer looking at reversed lines. In the EEG data, the difference between panels without lines and those with conventional lines tended to be larger among less experienced readers, while the difference between panels with reversed lines and those without lines grew with experience. The authors interpret speed lines as a learned expressive convention rather than a direct consequence of biological vision[21]. "Effect lines" is a broad umbrella term; this study specifically examined speed lines attached to moving people and objects. It did not establish that the findings extend to other effect lines, such as focus lines. Figure 2 was created to illustrate the effect of the lines and does not reproduce the study's stimuli.
Figure 2: With and without speed lines
This illustrative comparison duplicates an AI-generated image of a car and adds speed lines only to the left panel. The car on the left appears to be moving, while the one on the right appears stationary.
When people look at pictures, they rely both on perceptual tendencies that help them read facial expressions and on expressive conventions learned from the works they have encountered. We can therefore understand large eyes as a device that draws on our tendency to use facial changes as emotional cues, and speed lines as an expressive device whose interpretation is shaped by experience with manga. Highlights involve both the ease of distinguishing fine changes in the eyes and conventions learned from earlier works. Each device combines perception and learning in a different way.
How we interpret expression cannot be divided neatly into what is "innate" and what is "learned from culture," because development, experience, medium, and context all interact. The account in the main text does not come from experiments that tested large eyes or highlights in isolation; it applies research on face perception and manga speed lines to these artistic devices. Large eyes draw on tendencies in face perception, but which eye sizes and drawing styles feel natural also depends on familiarity with the style.
As large eyes and speed lines show, expression is not merely a one-way transfer of meaning from creator to audience. Works draw on the interpretive habits that audiences already possess while gradually changing those habits. If communication is not simply the delivery of a finished message, what does a work communicate, and how?
2. A work's meaning arises from its relationship with the audience
The simplest account is a transmission model of meaning: the author puts meaning into the work, and the viewer takes it out. This model works reasonably well for signs and explanatory text. In art, however, the same cues can give rise to different meanings and emotions depending on the surrounding context and the audience's experience. A work is not merely a container for finished meaning; it can also be a catalyst through which meaning and emotion arise in the audience.
The one-way model of "author → work → viewer" works well for clear messages such as signs, notices, and instructions. Applied to art as a whole, however, it fails to capture meanings discovered during the creative process, experiences that cannot be put into words, and individual viewers' associations. This discussion draws on reception theory, in which readers construct meaning by filling gaps in a text, and on arguments that works remain open to multiple readings[22, 23]. Figure 3 does not directly represent any one of these theories; it combines several perspectives, including the medium, the viewer's state, and cultural knowledge.
[23] Umberto Eco, "The Open Work," Harvard University Press (English translation, 1989), 1962. (The "open work" as controlled openness)
Figure 3: How meaning arises from a work
The work's cues, its presentation, the viewer's state, and accumulated experience and cultural knowledge combine to produce meaning and emotion. The orange arrow shows that the experience of encountering a work changes how the next work is seen.
Consider the face with its eye highlights removed. In isolation, it is simply a gloomy-looking face, but when shown immediately after a cheerful conversation, it reads as a sudden change in the character's inner state. A viewer familiar with the anime convention of manipulating highlights can also interpret the change in the eyes as a narrative cue. Viewers may interpret the same image differently depending on the surrounding scenes and the conventions they have learned.
Narrative context is not the only factor that shapes reception. A picture displayed on a museum wall creates a different experience from the same picture encountered while scrolling on a smartphone: its apparent size and surroundings differ, as does the time a viewer is likely to spend with it. How people experience a work depends not only on the work itself but also on where and how it is presented.
For creators, too, meaning is not necessarily fully formed in advance. They sometimes discover what they are making during the process by trying different materials, making use of an accidental bleed, or rereading what they have written[4]. Nor can audiences' responses to works always be fully captured in words. The sense of time created by music or the bodily sensation of standing before an abstract painting cannot be conveyed adequately by text alone. That is precisely why experiencing the work itself matters[5, 6].
That said, the existence of multiple readings does not make every reading equally valid. One can read a mystery novel as a cookbook, but the text offers almost no support for doing so. A work's form, word order, composition, medium, genre, and historical context shape the range of reasonable interpretations. Although the author's intent is not the final word on a work's meaning, it remains relevant evidence, as do the circumstances in which the work was created.
Some theories argue that readers give form to a work by filling in what remains unwritten; others emphasize that works admit multiple readings or that communal conventions determine which interpretations are valid[24]. Several theorists have also challenged the treatment of authorial intent as the final arbiter of meaning, though with different emphases[25]. All of these approaches give the audience an active role in the emergence of meaning. Rejecting authorial intent as the final arbiter, however, is not the same as declaring the author irrelevant. The circumstances of creation and the author's historical context remain relevant to interpretation.
[25] Roland Barthes, "The Death of the Author," Image―Music―Text (1977), 1967. (Separating the author's intent from the work's meaning)
Large eyes, highlights, and speed lines are all cues created by the artist. Artists cannot impose a single interpretation on their audiences; they choose which cues to provide and which perceptions and emotions to try to evoke. As the feedback arrow in Figure 3 shows, works not only draw on viewers' existing interpretive habits; through new compositions, symbols, editing techniques, and genre conventions, they can also change how later works are perceived[7, 8]. How, then, do these acquired habits take shape beyond the individual?
3. Culture cultivates new interpretive habits
When works change people's interpretive habits, the effects do not remain confined to individuals. As people discuss and imitate works, teach others how to interpret them, and apply those interpretations in new works, the habits come to be shared and are passed on as part of culture. Creative methods are passed on in much the same way, but here we focus on the interpretive habits that guide what people notice, what they feel, and how they anticipate what comes next.
This process extends beyond visual art to words and narrative conventions. From a single term such as "elf" or "villainess," readers familiar with fantasy can infer a character's appearance, the setting, and even plot developments that the work never spells out[9, 10]. When author and reader share background knowledge, the work need not explain everything; brief cues allow readers to fill in what remains unwritten. In this sense, culture provides the background knowledge that allows much to go unsaid. The following table illustrates elements that experienced readers may infer from a single term and ways an author might subvert those expectations.
What a term evokes varies by reader, era, and body of works; none of these associations is fixed. We draw on a corpus study of vocabulary and genre trends in web novels[9] and on research into how genre guides readers' expectations[10]. The specific entries in the table, however, are examples devised for this article. Huang's study did not measure associations with individual attributes, such as the proportion of readers who think "long-lived" when they encounter "elf," or test readers' plot predictions experimentally. Nor do these conventions cause every reader to derive the same meaning. Because audiences fill in what remains unwritten using their own knowledge, their associations differ. In cognitive science, structured knowledge about common situations is called a frame or script; in research on conversation, information that participants believe they share is called common ground[26, 27, 28]. Our use of background knowledge here applies these ideas to the reception of works.
[9] 黄 晨雯, "ビッグデータとしてのウェブ小説――言語特徴およびトレンド解析――," 大阪大学博士論文, 2022. (Large-scale analysis of linguistic features and trends in Japanese web novels)
[10] John Frow, "Genre (2nd ed.)," Routledge, 2015. (Genre as a framework directing readers' expectations)
[28] Herbert H. Clark, Susan E. Brennan, "Grounding in Communication," Perspectives on Socially Shared Cognition (APA), 127―149, 1991. (Common ground and calibrating how much to explain)
Term
What experienced readers tend to expect
Example that subverts the expectation
Elf
Pointed ears, a long life, a forest home, skill with a bow
An elf with an office job in the city
Villainess
Aristocratic society, a broken engagement, impending doom, an attempt to avert it
A protagonist who does not avert her downfall and instead chooses her own path after being condemned
Death flag
A character dying in battle just after saying, "Once this battle is over…"
A character raises the flag but survives anyway
These conventions do not merely constrain authors. Through repetition, the plot in which a villainess averts her doom has itself become a convention that readers expect. Once a pattern of foreshadowing acquires a name, as with the "death flag," readers can recognize how a story is shaping their expectations, and authors can play with that awareness. Deviations become meaningful precisely because the underlying expectations are shared. Once imitated, a subversion can become the next convention, setting up yet another round of subversion[10].
Figure 4: The cycle of art and culture
The diagram traces how works and interpretive conventions change each other and influence later creations through imitation, divergent interpretation, and repetition.
Works present a variety of cues, and audiences draw on existing interpretive conventions while sometimes acquiring new ones. People share these conventions, and later artists reuse them unchanged or recombine them into new forms. This cycle, in which works and interpretive conventions change each other and give rise to new creative work, is what this article calls "research and development in communication." Here, "communication" extends beyond conveying a meaning already fully formed in the creator's mind. It includes a work's capacity to evoke perceptions, emotions, associations, and questions in its audience. In this metaphor, "research" consists of using works to test what different cues evoke, while "development" consists of imitating, altering, combining, and reworking those cues into new forms of expression.
We use "communication" broadly enough to encompass works that resist paraphrase in a single sentence, such as music and abstract painting, as well as forms of expression that emerge through accidents in the creative process. In each case, the work changes the audience's perceptions and emotions, and that experience shapes how people interpret and make subsequent works. Even when someone creates solely for themselves and shows the work to no one, the cycle still applies once they respond to what has taken shape and adjust their approach. On the other hand, this view does not fully capture the value that resides in the act of making itself, apart from the reception of the finished work. The framework can likewise address how rituals create shared experiences among participants, but communication alone cannot explain how rituals fulfill obligations and maintain communal bonds.
"Research and development" here is not limited to corporate activity governed by plans and KPIs. Nor does it necessarily involve a single creator working toward a defined goal with a fixed method. Instead, refinements, cross-genre borrowing, technical constraints, and audience responses interact across creators, audiences, and generations, changing expression and interpretation together. Research and development are not separate, sequential stages either. Art also has values unrelated to novelty, including preservation, consolation, play, and the transmission of craft traditions. This view does not reduce the value of art to a single criterion; it is a way to focus on the moments when works and audiences change each other. Scholars continue to debate the extent to which human culture can be considered "cumulative"[29].
[29] Krist Vaesen, Wybo Houkes, "Is Human Culture Cumulative?," Current Anthropology 62(2), 218―238, 2021. (Methodological critique of the received view of cumulative culture)
This research and development is not a linear progression that prizes novelty above all else. It is a process of cultural trial and error that includes imitation, divergent interpretation, and repetition. As the interpretive habits that emerge from it are shared and passed on, culture becomes a repository of those habits.
This accumulation does not advance in a single direction, nor is it shared uniformly. What people associate with a work varies by region, generation, field of expertise, and community[11, 12]. Interpretive conventions change through criticism, translation, parody, divergent interpretation, and adaptation to other media; some spread widely, while others are forgotten. Researchers have found differences across languages and cultures even in something as basic as color categorization[13].
Shared interpretive conventions allow brief cues to convey a great deal, but to someone who lacks the necessary background, those cues may look like a mere lack of explanation. Nor do conventions become widely shared on their own. Education, criticism, markets, and streaming-platform recommendation systems all influence what counts as standard[14, 16]. Stereotypes about people and groups can spread through the same mechanism. The efficiency of leaving things unsaid and the danger of reinforcing prejudice and exclusion are two sides of the same mechanism.
Scholars have analyzed how everyday signs make particular ways of seeing appear natural and how classification systems render certain people less visible[30, 31]. The knowledge needed to understand works is distributed unequally across classes, regions, and levels of expertise and can function as cultural capital that advantages those who possess it. Education and recommendation systems can reduce this imbalance, but they can also entrench one way of seeing as the standard. Pictograms are not necessarily culturally neutral either. Using a skirted figure to represent women embeds the cultural assumption that links clothing to gender. Binary pictograms are also poorly suited to guiding some users or identifying some facilities.
[30] Roland Barthes, "Mythologies," Editions du Seuil, 1957. (Mythologies: the naturalization of ideology through cultural codes)
Whether a form of expression survives within a culture depends not only on its quality but also on institutions and systems of distribution. Moreover, quality can be measured in more than one way. To evaluate what happens when generative AI makes ideas easier to turn into drafts, we first need to clarify what "communicating well" means in each context.
4. What counts as communicating well depends on the purpose
Signage has a clear job to do, making it a useful example of how the criteria for evaluating communication vary by purpose. Effective communication is not determined by the amount of information or the degree of realism alone. The subway map that you consult in a station does not depict distances or track layouts accurately; it emphasizes station order and transfer points. By omitting some information and emphasizing other details, it focuses on what subway passengers need to know (Figure 5). It makes station order and transfers easy to identify but is poorly suited to judging walking distances or directions above ground.
Schematic transit maps that prioritize station order and transfers gained widespread recognition after the London Underground adopted Harry Beck's design in 1933[32, 33].
[32] Harry Beck, "Railway Map," London Museum Collection, 1935. (Harry Beck's 1935 route map, held by London Museum)
[33] Transport for London, "TfL History: Our Brand Assets," Made by TfL Blog, 2024. (The design of the earlier maps and Beck's map, published in 1933)
Schematic map
Geographic map
Figure 5: Schematic and geographic maps of central Tokyo
The map on the left is a square crop of central Tokyo, centered near Tokyo Station, from Tokyo Metro's simplified route map. The map on the right plots the stations of 8 Tokyo Metro lines, 4 Toei Subway lines, and the JR Yamanote Line at their geographic coordinates across roughly the same area. Although the two maps show different sets of lines, the comparison demonstrates how the schematic map reorganizes the layout to prioritize station order and transfers, substantially distorting actual distances and directions.
Omitting detail, however, does not automatically make a design effective. Signage relies on multiple cues beyond shape, including color, the arrangement of elements, where a sign is installed, and familiar configurations. Restroom pictograms, for example, combine differences in figure shape with color and a familiar side-by-side layout. Repeated exposure, placement, and standardization teach and reinforce these associations[15].
Figure 6: Clear cues to intended users (left) and an unclear facility mapping (right)
The design on the left combines differences in the human figures with color and a familiar composition. The design on the right is an illustrative reconstruction of two curved, abstract shapes. From this figure alone, it is difficult to infer which shape corresponds to which facility; communicating the distinction requires additional cues such as labels, placement, and repetition.
The subway map omits accurate distances and track layouts to help passengers identify station order and transfers. The restroom pictograms combine differences in figure shape with color and a familiar composition so that people can quickly infer the intended users. At the same time, they reflect a view that links gender to appearance and divides users and facilities into two categories. By contrast, a design centered on differences in shape may distinguish the figures without making clear which facility each one indicates.
Which cues to retain depends on whether the goal is rapid recognition, guiding a diverse range of users, or challenging the existing classification. The key is not simplicity itself, but what kinds of understanding, behavior, or experience the design should produce—and for whom. Signage is valued for using shared conventions to prompt quick action, but art can also preserve multiple interpretations or prolong the act of perception by creating a sense of estrangement.
One influential theory holds that art makes the familiar strange and slows the act of perception[8]. Communication can be judged by speed, accuracy, memorability, satisfaction, intensity of experience, or novelty, with the relevant criteria depending on its purpose. Advertising, architecture, film, and conceptual art each combine usefulness with some degree of openness to interpretation. Institutions such as museums and markets also shape what is treated as art. Rather than classifying an object once and for all as either art or design, it is clearer to ask what purpose matters in the evaluation at hand.
[8] Viktor Shklovsky, "Art as Device," Russian Formalist Criticism: Four Essays (University of Nebraska Press, 1965), 1917. (Defamiliarization: art's work of slowing automatic perception)
Even after we decide what effective communication means, technology can change both the means available to us and which parts of expression require the most effort. Photography marked one such major turning point.
5. When technology changes, creative challenges change too
Realistic depiction does not arise naturally either. Creating a sense of depth on a flat surface requires techniques for handling perspective, shading, contour, and color; viewers likewise learn to infer depth from lines, light, and shadow. Linear perspective is one such way of communicating, developed over a long period.
Linear perspective is a convention systematized in a particular time and place, not the universal origin of pictorial representation. Other traditions of representing space coexist with it, including East Asian landscape painting. The cited art-historical study argues that representation develops as artists repeatedly correct existing schemata through observation and that viewers, too, learn how to see pictures[34].
Photography changed the effort and cost required to create and reproduce accurate visual records. As paper-based photography spread in the 1850s, a single negative could produce many prints, making visual records easier to reproduce and distribute than hand-painted portraits[17]. At the same time, new skills emerged around taking and developing photographs, and photography became a craft in its own right. The change is better understood not as painters' work simply disappearing, but as a shift in the uses of painting and photography, the skills they required, and the qualities valued in each.
Paper photography made it possible to produce dozens and sometimes hundreds of nearly identical prints from a single negative. As demand for photographs grew, studios multiplied, public documentation projects began recording architecture and disasters, and the range of both image makers and uses expanded. At the same time, early paper photography was labor-intensive: photographers adapted chemical formulas to conditions such as temperature and shaped texture and tone through exposure and the processing of negatives and prints[17]. Although it was a technology of reproduction, it still demanded individual mastery of materials and processes.
Nor does this mean that the value of painting before photography lay solely in recording appearances accurately. In nineteenth-century France, for example, the state-sponsored Salon and the Academy were central institutions in the evaluation of painting, with a formal hierarchy of genres headed by history painting[35]. This example does not represent every form and use of pre-photographic painting worldwide, but it shows that accurate depiction was not the only criterion of evaluation.
Walter Benjamin also famously argued that technologies of reproduction change the meaning of a work's uniqueness and the aura associated with it[36]. Similar comparisons can be drawn between recorded music and live performance or between letterpress printing and hand copying: older skills do not necessarily disappear, but they may be put to new purposes or survive because the process of making is valued in itself.
[17] Malcolm Daniel, "The Rise of Paper Photography in 1850s France," Heilbrunn Timeline of Art History, 2008. (Reproduction technology, materials, and expanding uses in early paper photography)
A viewer standing before a work of contemporary art may be unsure what to look for. Contemporary art often explores concerns beyond realistic representation, and part of this history goes back to the spread of photography. By enabling accurate visual recording outside painting, photography changed the role of realism within painting itself. The same period also brought changes in painting materials, transportation, urban life, and exhibition institutions. Against this backdrop, Impressionist painters explored light and fleeting appearances; Cubists recomposed a single subject from multiple viewpoints; and the readymade presented manufactured objects as works. These are examples of how art broadened the questions it could explore. Realistic depiction continued, but as technological and social changes accumulated, art's expressive vocabulary and the range of qualities for which a work could be valued both expanded.
Artists now grouped as Impressionists depicted light, fleeting appearances, and modern life in different ways[37]. Cubism recomposed a single subject from multiple viewpoints[38], while the readymade presented manufactured objects as works, turning selection and the idea itself into artistic subjects[39]. Photography was only one factor: changes in painting materials, transportation, urban life, and exhibition institutions also contributed, so no single cause can explain painting's transformation. These movements overlapped chronologically, realist painting continued alongside them, and developments varied across regions and exhibition systems. We omit abstract painting from these examples because it cannot be reduced to a single question and its concerns overlap with those already discussed. Nor do we treat readymade artists as direct forerunners of generative AI users.
[37] Margaret Samu, "Impressionism: Art and Modernity," Heilbrunn Timeline of Art History, 2004. (The multiple factors involved in the emergence of Impressionism)
[38] Sabine Rewald, "Cubism," Heilbrunn Timeline of Art History, 2004. (Cubism's multiple viewpoints and recomposition of the picture plane)
New technology can change both what is difficult to make and what is valued in a work. Photography shows how lowering the cost of a particular task can shift the creative challenge elsewhere. Generative AI can likewise turn many ideas into drafts in a short time. What does that shift mean for creative work?
6. Generative AI through the lens of research and development in communication
Just as photography changed the effort involved in creating and distributing visual records, generative AI primarily changes the effort required to turn ideas into drafts. When this step becomes easier, the parts of the creative process that require time and judgment also change.
By "turning an idea into a draft," we mean turning a notion in one's head into a working version—a text, image, or structure—that can be evaluated and revised. Generative AI mainly reduces the time and labor required to do so. It does not eliminate the work involved in fact-checking, revision, rights clearance, or skill development, nor does it eliminate the cost of electricity and computing resources. The magnitude of the change also varies by task, user, and model. Moreover, whereas a photograph is normally formed by light arriving from its subject, a generated image does not provide evidence that the event it depicts actually occurred. Generative AI raises distinct problems, including questions of consent and compensation for training data, imitation of artistic styles, and the distribution of profits. The comparison here is limited to one shared feature: both technologies reduce the labor involved in a particular stage.
To understand this change, we divide the creative process into five stages that makers revisit as needed: finding a question or problem; conducting research and defining constraints; turning ideas into drafts; selecting, refining, and verifying; and completing and delivering the work. Generative AI can contribute to any of these stages (Figure 7).
This framework draws on the 24-item Creative Process Assessment Scale (CPAS), which measures creative activity across eight phases. This article consolidates those phases into five to make the impact of generative AI easier to see[40]. The CPAS study examined the scale's structure using responses from 2,324 people, but the five stages used here are not a direct reproduction of that scale. Nor are they meant to prescribe the order in which creative work proceeds; in practice, people move back and forth among the stages, and different people and tools may handle each one.
Figure 7: The creative process and generative AI
The diagram divides the process from finding a question to delivering the work into five stages. In practice, people move back and forth among them, and different people and tools may handle each one. The blue background band represents the guiding aim—what to convey and what experience to create—that spans the entire process. Generative AI most directly reduces the effort and cost of "turning ideas into drafts," shown in orange. Audience reception of the completed work leads to the next question.
Even if many ideas can be turned into drafts quickly, the time available to compare them, choose those that fit the purpose, check them for errors, and polish them does not increase accordingly. If only drafting becomes easier, the main difficulty shifts to selection, revision, and verification.
AI can also support selection, revision, and verification. Such support changes where time and judgment are needed, though the burden falls on different stages depending on the task, user, organization, and tool.
One framework distinguishes three forms of creativity: combining known elements, exploring a space of possibilities, and transforming that space itself. AI can contribute to all three[41]. This does not mean, however, that AI can determine goals and values on its own; the people and organizations involved in the creative process must decide which proposals and criteria to adopt. Moreover, editing according to criteria chosen by the editor differs from checking output against criteria set in advance. Even if AI shifts work from making to choosing, workers' judgment and skills do not automatically improve, nor does their work necessarily become more meaningful[42]. The cited paper offers a conceptual, ethical analysis of generative AI and the meaning of creative labor; it does not empirically measure the magnitude of its effects on workers.
[41] Margaret A. Boden, "Creativity and Artificial Intelligence," Artificial Intelligence 103(1―2), 347―356, 1998. (Computational creativity as combination, exploration, and transformation)
The impact of generative AI also depends on how we measure it. Its effects cannot be captured by a single yardstick. AI can speed up work and sometimes improve evaluators' ratings of individual works, but outputs may also become more similar to one another, and examples shown to users can anchor their ideas. Production speed, the quality of an individual work, and the diversity of a set of works must be considered separately.
In a randomized experiment with 453 participants, access to ChatGPT reduced the time participants spent on professional writing tasks by 40% on average and increased evaluators' quality scores by 18%[43]. The study involved short writing tasks completed by college-educated professionals; it did not examine longer-term creative work, factual accuracy, or skill retention. In another randomized experiment, 293 people wrote eight-sentence stories. Participants who had scored low on an earlier divergent association task (DAT) saw the largest gains in novelty and usefulness ratings when they used ideas from GPT-4. At the same time, stories written with GPT-4's ideas became more similar to one another[44]. Completion time, evaluator-rated quality, novelty, number of ideas, differences among works, and user satisfaction are distinct measures; an improvement in one does not guarantee improvements in the others or the preservation of long-term skills.
In an experiment in which 60 people designed chatbot avatars, participants in the AI image-generation condition showed greater "fixation" on the initial example and produced fewer, less varied, and less original ideas than the control group[45]. A study comparing 22 LLMs and 102 people on three standardized verbal creativity tasks found that the LLMs' responses were more similar to one another than the human responses were[46]. This was a group-level comparison between models and people working independently on verbal tasks; the study did not directly examine human–AI co-creation, images, or music. A meta-analysis of 28 studies involving 8,214 participants likewise reports higher creativity ratings with AI support alongside lower diversity of ideas[47]. However, that manuscript has not yet been peer-reviewed, its synthesis of diversity effects covers only 4 studies, some of the included studies are preprints, and a single researcher performed most of the study selection.
By contrast, an experiment in which 10 AI personas with different backgrounds and creative orientations generated story ideas suggested that the diversity of their outputs could match that of work produced entirely by humans[48]. However, this was a single experiment with limited statistical power in its second stage, so it does not show that homogenization can always be prevented. The degree of homogenization may depend not only on the model but also on design choices, such as whether every user receives the same support or the system incorporates different perspectives. A study using submission records for more than 4 million works found that after AI adoption, average content novelty and visual novelty declined; maximum content novelty rose slightly, while maximum visual novelty fell[49]. Because this was a difference-in-differences analysis of observational data and users chose for themselves whether to adopt AI, the findings can be interpreted causally only if the method's assumptions hold.
[45] Samangi Wadinambiarachchi, Ryan M. Kelly, Saumya Pareek, Qiushi Zhou, Eduardo Velloso, "The Effects of Generative AI on Design Fixation and Divergent Thinking," Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 1―18, 2024. (Experiment with 60 participants on design fixation from viewing AI-generated images)
[46] Emily Wenger, Yoed N. Kenett, "Large Language Models Are Homogeneously Creative," PNAS Nexus 5(3), pgag042, 2026. (Diversity of responses by 22 LLMs and 102 people on standardized verbal creativity tasks)
No metric automatically takes priority. Even when AI proposes evaluation criteria, someone must decide which ones to adopt, how thoroughly to verify the results, and what to deliver to whom. The people and organizations involved in creating the work must take responsibility for these judgments and their consequences.
Even after a creator chooses a promising draft and completes it, the time audiences can spend with works does not increase. If the volume of generated content in circulation grows while audience attention remains limited, each work will tend to become harder to discover and less likely to attract sustained attention. Consequently, reception is likely to depend not only on a work's appearance but also on its provenance—who made it and how—and on the channels and intermediaries through which it reaches its audience.
Among the works audiences encounter, those they discuss, imitate, and preserve help shape how future works are interpreted. What survives depends not only on a work's quality but also on audience attention, criticism, education, distribution, recommendation and preservation systems, funding, and chance. The ability to turn ideas into drafts does not by itself determine how this cycle develops.
Herbert Simon observed that as information becomes abundant, recipients' attention becomes scarce[50]. Trust, relationships, and venues for showing work are likewise limited, and recommendation algorithms strongly shape who encounters what. Moreover, labeling a work "AI-made" or describing its production process can change how people evaluate it, even when the work looks the same[51]. The Coalition for Content Provenance and Authenticity (C2PA) has developed a technical specification for binding tamper-evident records of creation and editing history to digital assets such as images. The public-facing term "Content Credentials" refers both to the technology as a whole and to the C2PA Manifest bound to each asset[52]. The specification can verify, for example, that a manifest is bound to a particular asset and has not been altered since it was signed. It does not guarantee that the manifest's contents are true, that the recorded history is complete, or that the work is of high quality. Unrecorded steps are simply absent, and the lack of a record does not prove that a work is fake. Likewise, quality alone does not determine whether a work endures as part of a culture. Inclusion in textbooks, prizes, recommendation algorithms on platforms, preservation budgets, rights holders' policies, the availability of translations, and chance events such as disasters or the dispersal of collections have all shaped what remains.
[50] Herbert A. Simon, "Designing Organizations for an Information-Rich World," Computers, Communications, and the Public Interest (Johns Hopkins Press), 37―72, 1971. (The abundance of information and the scarcity of recipients' time and attention)
Anime's large eyes, the disappearance of eye highlights, and manga speed lines form part of an expressive vocabulary that draws on both human perception and interpretive conventions learned through encounters with works. Works do not merely carry a fixed, finished meaning; they evoke meaning and emotion in their audiences and, in doing so, shape how later works are interpreted. This cycle, in which expression and interpretation change each other, is what this article calls "research and development in communication."
Generative AI is making it easier to put new works into circulation, especially by making ideas easier to turn into drafts. Yet even as our capacity to create grows, it does not automatically determine what to select and refine, whom to reach, or what aspects of the works audiences receive will endure as part of the culture. In the creative process, selection, revision, and verification take on more importance; in distribution, audience attention and trust can become even greater constraints.
The production of this article reflects that change. Generative AI was used to research the topic, explore examples and counterarguments, design the structure, draft the text and figures, and edit the manuscript files. The human did not edit the manuscript directly; instead, through dialogue, they set the theme, acceptance criteria, and publication policy and directed the production process.
For this article, the human set the theme and central claim—framing art as research and development in communication—while generative AI proposed sources, counterexamples, alternative structures, passages, and figures. Through dialogue, the human chose among these options, checked the strength of the claims, the scope of the citations, and the purpose of each figure, and directed revisions. The AI contributed not only to the "turning ideas into drafts" stage in Figure 7 but also to research, evaluation, and verification. The human retained responsibility for deciding which proposals to adopt and what to publish. This is one possible production process; it does not demonstrate that the method generally improves article quality.
Now that so much more can be made, what to select, what kinds of experiences to create, and whom to reach matter more than ever. Could grappling with that challenge provide a starting point for rethinking creative work in the age of generative AI?
[4] John Dewey, "Art as Experience," Minton, Balch and Company, 1934. (The interplay of doing and undergoing in making; art as experience moving toward completion)
[5] Susanne K. Langer, "Feeling and Form: A Theory of Art," Charles Scribner's Sons, 1953. (Art as a symbol presenting forms of feeling that cannot be stated as propositions)
[7] Nelson Goodman, "Ways of Worldmaking," Hackett Publishing, 1978. (Recomposing the world's classifications through symbol systems)
[8] Viktor Shklovsky, "Art as Device," Russian Formalist Criticism: Four Essays (University of Nebraska Press, 1965), 1917. (Defamiliarization: art's work of slowing automatic perception)
[9] 黄 晨雯, "ビッグデータとしてのウェブ小説――言語特徴およびトレンド解析――," 大阪大学博士論文, 2022. (Large-scale analysis of linguistic features and trends in Japanese web novels)
[10] John Frow, "Genre (2nd ed.)," Routledge, 2015. (Genre as a framework directing readers' expectations)
[11] Umberto Eco, "A Theory of Semiotics," Indiana University Press, 1976. (Cultural knowledge as an encyclopedia rather than a dictionary)
[14] Pierre Bourdieu, "Distinction: A Social Critique of the Judgement of Taste," Harvard University Press (English translation, 1984), 1979. (Aesthetic judgment and cultural capital; the unequal distribution of interpretive conventions)
[17] Malcolm Daniel, "The Rise of Paper Photography in 1850s France," Heilbrunn Timeline of Art History, 2008. (Reproduction technology, materials, and expanding uses in early paper photography)