top of page
Blurry Blue_edited_edited.jpg

Much ado about “Good”: Evaluating multilingual creative quality

Updated: Jul 16

In the previous quality articles, I argued that localization quality is broader than linguistic correctness. It is shaped by purpose, audience and the expectations negotiated between the requester and the provider. Depending on the content, it may also need to account for the user experience and the product journey. But what happens when copy should also make someone stop scrolling, smile, remember a brand or feel something? This can happen across many content types. This article focuses on one in particular: creative content, where “good” becomes an even more opinionated subjective concept.

As LLM output continues to “improve", experimentation with AI-generated creative copy grows under the promise of hyperlocalized content. Whether it should be used, where and under which conditions is a discussion for another day. But as its use grows, so does a familiar quality management question: how do we measure something that has no single correct answer, no source text and, sometimes, only a brief saying "make it engaging"?

Creative assets make the quality question especially interesting because they sit at the intersection of language, culture, emotion, brand and design. A sentence may sound natural but forgettable, culturally fluent but off-brand, or correctly written and wrong for the image next to it. So, once again, let me try to quantify the unmeasurable: what makes multilingual creative content “good”, how can we evaluate it with something more rigorous than gut feeling, and what does the resulting framework tell us about the context an AI system needs.

Define before you evaluate

If you attended the first AI Localization Think Tank Conference, AI ThoughtCon, in March 2026, you may remember that we introduced DEM: Define, Evaluate, Measure. DEM provides a structured way to approach AI localization quality, when there is no consensus on what “good” means. First, define the outcome and the risks that matter. Then decide how to evaluate them. Lastly, turn the results into a score or metric.

The Define stage can sound as if we were looking for one correct answer against which every output can be compared. However, most of the time we are really agreeing on purpose, expectations, and acceptable risk —or, borrowing Agustín Da Fieno Delucchi's terminology, defining the floor of quality. What does this content need to achieve? What cannot go wrong? Which, at the same time, means that we do not tend to define its full potential or ceiling. The brief may provide direction, but the goal is to create the intended effect for an audience, brand, and market… leaving plenty of room for individual interpretation.

Before going any further, it is worth clarifying what I mean by multilingual creative content. I am not talking about novels, poetry, or literary translation (that would require a different article and a different author). I am talking about marketing assets: display ads and banners, social media posts and carousels, hero banners and promotional copy, print and out-of-home advertising, paid search ads, push notifications, and in-app messages. The kind of content that uses very few words and still generates a surprising number of opinions and, when wrongly rendered, may damage a company's reputation.

From usability to resonance

Experience deserves a broader definition. When I said quality had to meet the intended experience, that could easily be interpreted as user experience: usability, friction, and whether someone can successfully complete a product journey. That is not wrong, but it is not the full concept. In creative content, experience extends beyond usability: it aims to create engagement and resonance. It is content trying to be felt.

Usability asks, “Can the user complete the task?” Creative content asks, “Did it do the trick?” Yes, they are different questions, although stakeholders expect both to be answered at once when generating AI marketing assets. That distinction matters because traditional localization quality frameworks evaluate usability well but rarely assess the target purpose and geolanding.

The missing layer is creative effectiveness.

A modular framework for creative quality

Defining "good" in creative localization means first deciding which layer of quality actually matters. Once those layers are defined, we can evaluate each one independently before deciding whether, and how, to aggregate them into an overall quality view.

  • Language quality*

  • Creative effectiveness

  • Context coherence

  • Design quality

*AI-oriented taxonomy beyond the traditional localization quality management categories, including hallucination, safety, bias, and appropriateness.

This framework helps experts evaluate AI-produced assets consistently during the human assessment step. Whenever possible, those evaluations should be inter-annotated, correlated, and validated against other evidence or metrics like user research, CSAT, hook rate, engagement, or conversion. Expert judgement tells us whether content appears likely to work. User behaviour tells us whether it actually did.

This layered approach is also reflected in MMQEval (Da Fieno, Gašová and Soukeník, 2025), a recent framework for multimodal evaluation. MMQEval groups quality dimensions comparable to the above, and evaluates them using scalar ratings. I believe these similarities are encouraging (in fact, reassuring) because they suggest a shared instinct for what we need and what works. My objective, however, is slightly different: to provide localization teams with a framework that is comprehensive enough to capture creative quality, while remaining simple enough to implement it on Monday.

Much Ado About Nothing, by Alfred W. Elmore (1846). Oil on canvas.
Much Ado About Nothing, by Alfred W. Elmore (1846). Oil on canvas.

Creative effectiveness

Creative effectiveness is where the framework asks for the very first time a difficult question: "Is this good?" Unlike terminology, there is no source against which creative effectiveness can be checked. So we have to define "good" on its own terms and break that judgment into smaller categories that can be evaluated. Those are:

  • Creative craft: how effectively the copy uses creative writing techniques within the target language: hook, emotional impact, rhythm, etc.

  • Cultural relevance: whether those techniques are appropriate and inclusive for an audience and market: references, humour, metaphors, imagery, symbolism, etc.

If you have one and not the other, your creative effectiveness might still be a fail, and scoring them together under one category would hide that gap; scoring them separately makes that gap visible and ready to be addressed.

Existing frameworks like MQM and DQF already recognize cultural quality through dimensions such as Audience Appropriateness and Locale Conventions. What has been missing is evaluating those dimensions with the same rigour applied to Accuracy, for example. The difference with this framework is that rather than treating Creative effectiveness as a defect category, I propose evaluating it using a weighted scale, aligned with MMQEval.

5 – Crafted: memorable and resonant.

4 – Engaging: written with purpose.

3 – Does the job: competent and on-brief.

2– Flat: functional yet mechanical.

1 – Alienating: stereotypical and offensive.

One thing worth saying before anyone starts treating this as a leaderboard is that a 5 is not always the right answer, and a 2 is not always a bad copy. Creative effectiveness only has meaning in relation to a brand strategy and an asset type. For example, a transactional message or a legal disclaimer was never briefed to become memorable. Judging those assets as “crafted” just because they sound good would be grading them against a brief nobody wrote. The scale measures whether the content achieved the intended creative outcome, not whether it was the most creative piece of copy in the room. So, sorry, no benchmark. Unless it is completely off; that is why we have an “Alienating” category as an immediate fail gate.

AI's limitations become more visible here. LLMs may naturally lean towards patterns that look generically engaging. That may produce fluent copy, but fluent is not the same as culturally resonant. Unless we explicitly provide the model with the brand voice, creative constraints, and cultural context, it has little reason not to optimize towards the statistical average of "good marketing copy." Fortunately, this framework also tells us what information production systems need to produce better output.

Context coherence and multimodal fit

Creative assets distribute meaning across text and visuals. A sentence reviewed on its own can be linguistically correct and still be wrong the moment you look at the image.

Unlike MMQEval, which evaluates Intra- and Cross-modal Consistency on a scalar scale, this framework treats Context coherence as a separate pass/fail layer. It sits outside Creative effectiveness because alignment between text and visual is not something that is "more or less good"; either it exists or it does not.

Typical visual-textual mismatches include:

  • gender or people, e.g. refers to a woman and the visual shows a man;

  • products or objects, e.g. promotes a laptop while the image shows a tablet;

  • actions being performed, e.g. says “tap” when the image shows “swipe”;

  • season, location or setting, e.g. describes summer while it depicts winter;

  • cultural markers or symbolism, e.g. mentions Lunar New Year with Christmas decorations;

  • emotional tone between copy and visual, e.g. shows positive emotions while the visual shows disappointment.

Not every asset is suitable for AI

Drawing on Edward Hall’s distinction between low- and high-context communication, we know that not every creative asset depends on context to the same degree. On the one hand are relatively low-context assets: short, self-contained pieces of copy whose meaning is explicit with minimal reliance on cultural inference or visual-text coherence. On the other are highly context-dependent assets, like a social media video which may distribute meaning across visuals, timing, references, humour, and platform conventions.

Context dependency predicts two things. First, the amount of human effort required to review and adapt the asset. Second, the likelihood that an AI system will miss something important. The more meaning depends on visual, or situational context, the less useful it becomes to evaluate the text in isolation. Human reviewers reconstruct missing context naturally, while AI systems only do it if we explicitly provide it.

In other words, context is not simply another input to creative generation. It is part of the quality problem itself, and, as we all know, we are still trying to solve context in the age of AI.

Beyond Measure: Orchestrating Quality

The same framework that evaluates creative quality can also structure how we produce it.

Whether you use agents, RAG, automated checks, or human review, we should not bundle all quality layers of information together into a hierarchical, never-ending prompt that is supposed to hold a brief, a style guide, a brand voice, and a sense of visual composition simultaneously and get all four right at once. That is wishful thinking.

This is not an article on architecture. Still we need to stop treating "good" as an instruction handed to a model, or agreed with a stakeholder in an office corridor. Every guardrail written for an AI system is a definition already fought to make it precise during the Define stage. And it is also the other way around: writing the guardrail is what forces the quality definition to be precise in the first place.

This aligns with Da Fieno’s idea of “Context as a System”. Once quality dimensions have been made explicit, they stop being instructions hidden in a prompt and become persistent context for the system itself. Creative effectiveness and Context coherence become part of the environment in which a model operates, rather than something it is asked to remember. That matters specifically in hyperlocalization, where the same campaign brief requires different creative solutions from one market to another, and generic guardrails risk smoothing away the cultural differences they are meant to preserve. Skip all this, and you will move the prompt's ambiguity straight into quality that isn't good enough, which is harder to identify and even harder to fix.

Modular evaluation, modular production. The same principle applies to both ends.

Perhaps there was never much ado about measuring "good". The hardest part has always been agreeing on what "good" means and what it should achieve. If you are generating multilingual creative assets with LLMs, those quality layers become the context your system needs first.

 
 
bottom of page