top of page
Blurry Blue_edited_edited.jpg

“Build or Buy” your customized AI models: Roadmap to governance and transparency

1 day ago
7 min read

By Belén Agulló García

This is a short reflection article based on my +15 years of industry experience. It is not meant to be a thorough technical explainer, but rather a personal story about anguish and learning curves to help spark a conversation.

We need to use AI: what do we do?

Many localization professionals, whether working at an enterprise in a loc program or at an LSP, have at some point asked themselves this question, often full of anguish. Before LLMs, it was okay to delegate the customization of neural machine translation engines and the like to specialized providers, and it was okay not to know what was going on behind the scenes, as long as the expectations were met: MT output with decent quality that could be operationalized in post-editing workflows without risks. Or even raw MT that was fine for your use case.

With the advent of GenAI in 2022, the complexity for the localization teams increased, and so did the level of anguish. Now, everyone was running to implement AI in their workflows; there was pressure everywhere, loads of hype, promises, and expectations, but very few people actually knew what they were doing. If your company had an in-house machine learning or engineering team, you were privileged. You could work with them, learn from them, try to figure out what to do, create super cool solutions before anyone else, and signal innovation. You became the cool kid on the block while the rest of us mortals were feeling clueless, trying to catch up.

Image: Alicia Silverstone looking clueless while wearing her iconic yellow outfit in the movie Clueless (1995).
Image: Alicia Silverstone looking clueless while wearing her iconic yellow outfit in the movie Clueless (1995).

But if you didn’t have access to these expensive and highly specialized resources, then you had two options: you could either take a crash course on how to code and how to fine-tune and LoRA and prompt and RAG and [insert LLM-related jargon], or you could go and find specialized companies that could do that for you. Some took the adventurous first path and succeeded, but again, not everyone had the luxury of time (or, let’s be honest, the motivation) to do it. Some preferred to delegate to others who actually knew what they were doing. This second option was safer and quicker to implement in the short term, but also riskier in the long term. We got the false impression that everything was under control, and that we didn’t have to learn all that technical stuff anymore. But that was a trap.

The trap: Believing that you can get by without learning the math

Okay, I must admit I’m not a big fan of maths and all the linear algebra (hello, embeddings) that goes into an LLM. Believe it or not, AI is not my passion, lol. So I was forced to learn all the technicalities behind LLMs. But that actually turned out to be one of my superpowers. Thanks to that learning effort, I can actually understand what happens behind the scenes when a tech provider says that their AI “learns on the fly” (their system uses RAG), or that their AI workflow selects the best engine for my translation (their system has a lookup table that routes certain content types to specific models based on a previously set up evaluation harness). I also know quality estimation models aren’t magical and that, most of the time, they need to be fine-tuned with specific data to be reliable, or that overfitting can bias demo results.

If you don’t possess a certain level of knowledge of how the tech actually works, you might end up with a failed solution that can’t deliver the results you expect. You might also become frustrated because you don’t know exactly how to fix it, and you have no choice but to trust your vendor. But we cannot afford not to know at this point. And the reason is that governance is now a big part of what localization programs and LSPs have under their purview. Someone once said, “You cannot manage what you cannot measure.” Now I say: you cannot govern what you don’t understand. You need to understand all the moving pieces, and where your team and expertise add value and can help prevent or mitigate risks in an increasingly automated localization pipeline.

Roadmap to governance and transparency

Okay, so now that I've convinced you to learn the math, what do you do? It all starts by having a proper AI strategy for your program (see here to learn more). In summary, you need to know what you want to achieve, how you will measure whether you’re achieving it or not, what your data strategy is, and how you will execute your AI master plan, technically speaking. Clear foundations are key to succeeding in the long term. But what does governance mean in a localization program? Some examples:

  • Content tiering matrix: Having a very well-documented content tiering matrix designed with criteria that match your localization requirements (cost, volume, languages, number of bugs, quality scores, turnaround times, etc.) and your business requirements (revenue, growth opportunity, engagement, user sentiment, brand reputation, content lifecycle, etc.). Of course, you cannot have an endless list of requirements, because at some point, it becomes impossible to operationalize. Your team needs to identify the few criteria that actually drive decisions, rather than trying to optimize everything at once. That’s your strategic advantage: what’s the most pressing aspect for your business right now, and how can you make a difference?

  • Localization content pipelines: The next step is defining your localization content pipelines and how each content type matches your pre-defined workflows, such as translation + editing, machine translation + post-editing, machine translation + quality estimation + AI post-editing / human post-editing, and so on.

  • Models: This is where you need to understand the math. Now that we’re here, you need to know which models you want to use for which use cases and why. How are you serving those models to your pipelines? And what level of customization (if any) is required? What does customization actually mean in your context? Is your orchestrated workflow using RAG, and how? Or perhaps your team performed low-rank adaptation on open-source models for error detection? Who is measuring quality, and how? And not only that, models are short-lived. They get deprecated once newer ones arrive. You have to constantly monitor models and switch to new versions when it makes sense; probably update your prompts, manage quota, limits, filters, and more.

  • Preparation for AI workflows: Once you know which pipeline and model you will use, what do you need to do to prepare your content to be successfully localized? Do you have all your curated linguistic assets in place? Your glossaries, your style guides, your annotated error data? Are they prepared to be effectively ingested by AI, and how? And who owns that preparation within your organization?

  • Guardrails: You need to identify which guardrails you need in place to avoid potential brand or legal risks before something bad actually happens in an AI workflow. Are there any geopolitical topics that are specifically sensitive to your brand? Or perhaps there is some terminology related to third-party IP that you need to follow, or otherwise you would be failing compliance? Identify the risks and put guardrails in place, such as additional AI checks or channeling certain content types or strings to in-country reviewers, to avoid any backlash. Here is where your localization and cultural experience shines, as a generic AI/engineering team may not have the same expertise.

  • Blindspots: Ideate a way to find blind spots in your strategy. Anticipate areas that your tracked data cannot fully cover and put a system in place to capture that less obvious information. For example, how will you know if a workflow is currently overkill for a certain content type or, on the contrary, the automation in place is negatively impacting certain aspects of your business?

Of course you don’t need to know every single aspect of the entire operation in depth; that’s why you work in a team. But you need to know enough of everything to make sure you’re controlling your AI strategy and not the other way round. A good analogy would be the role of a Product Manager. They need enough technical literacy to be a credible, effective partner to the development team, but they don’t need to write the code or architect the system themselves. The same goes for your localization program. You need to know how to talk to your machine learning teams or your technology partners to guide them in meeting the goals that your AI strategy is set to achieve.

So should I build or should I buy

You do you. I don’t think there’s a clear answer to this question. It depends on your circumstances, mainly your budget and time. Do what makes the most sense for your program and gets you closer to your goals. Building customized AI models internally gives you much more control over your data, quality, and workflows, but it comes with overhead costs that not everybody can afford. Buying customized models from an external vendor can give you access to expertise that you otherwise couldn’t get internally, although you might lose control over your data and be locked in if, at some point, you’re unhappy with their services or pricing.

And what if there was a third hybrid option? For some organizations, the realistic answer may be to buy foundational capabilities but keep orchestration, evaluation, data, guardrails, and some customization under internal control. That could also be an option worth exploring for your program.

It doesn’t matter which path you pick: you still need to learn the math. You still need to understand what’s happening under the hood so that you can fully govern your program and make it as impactful as possible.

Author’s note: This article was written entirely by a human. Any similarity to AI writing is purely coincidental. The infographic was created with Claude, though. Special thanks to Marina Pantcheva and Miguel Sepulveda for their great feedback on the article.

bottom of page