This year, three members of our team had papers accepted at EAMT, the annual European Association for Machine Translation conference that brings together researchers, practitioners, and industry professionals to share knowledge and discuss the state of the field. Between presenting our own work and attending sessions across the full program, we came away with a pretty clear picture of where MT research stands in 2026, and where it still has a long way to go.

 

What We Learned at EAMT 2026

Marina Sánchez Torrón, Ph.D., Senior Linguistic Engineer at Smartling, taking the stage at EAMT 2026.

 

Here are our biggest takeaways:

 

The debate over LLMs is over. Now comes the hard part.

Nobody at EAMT was asking whether to use LLMs for translation anymore. That question is settled. The conversation has fully shifted to how to deploy them well, and that's where things get interesting.

One of the strongest signals from the conference is that standalone LLMs are rarely sufficient for production environments. The emerging architecture isn't "LLM instead of translation memory." It's LLM plus translation memory plus retrieval. Translation memories aren't going anywhere. They're being repurposed as curated and vetted knowledge sources that inject company-specific linguistic knowledge directly into LLM workflows. High quality, clean TMs are increasingly the difference-maker, whether you're fine-tuning or building a RAG architecture.

Multi-agent translation approaches got some attention, but the honest takeaway is that they're still experimental. Jury's out on whether committee-style generation is production-ready anytime soon.

 

Terminology and style: progress, but not solved

If there was one theme that cut across almost every session, it was this: translating general language is largely easy. Translating specialized content consistently is not.

Terminology management remains one of the most significant unresolved production challenges in the field. A huge portion of the technical program focused on terminology translation, biomedical content, legal translation, and domain-specific adaptation, and the frustration was palpable. Style guide adherence came up again and again as a shared pain point, with no one claiming to have cracked it.

Researchers showed that LLMs have clear, measurable stylistic fingerprints compared to human translators: higher lexical density, more sophisticated vocabulary, a tendency to overuse dashes and underuse ellipses, and distinctive sentence-length patterns. Attempts to prompt LLMs to imitate specific literary translators mostly fell flat, with model-family effects (GPT vs. Claude) stronger than the style prompting itself.

A key conclusion: LLMs have made style measurable, but they haven't solved it. Human translators remain in a clearly distinct stylistic space that modern models can approach but not fully reach. AI post-editing can improve style and terminology compliance, but models frequently over-edit and human feedback remains essential for getting outputs where they need to be.

 

Fine-tuning still matters

Not long ago, there was a lot of enthusiasm for the idea that prompting would replace model adaptation. The conference evidence doesn't support that. Multiple papers showed clear benefits from domain-specific fine-tuning in legal, biomedical, literary, and technical domains. A LoRA-adapted EuroLLM-22B outperformed Claude Sonnet zero-shot on EU legislative translation.

The broader point: for use cases where terminology precision, style requirements, or regulatory constraints are critical, specialized adaptation remains valuable. Prompting and retrieval are increasingly important tools, but they don't replace domain expertise baked into a model.

 

Prompting has more nuance than most assume

Our paper on manual vs. automated prompt optimization (DSPy and GEPA) generated a lot of curiosity at our poster session. Most attendees hadn’t heard of GEPA, and many immediately saw its relevance to their own scaling challenges. The conversation didn’t stop at the hour mark and we were there through the coffee break. Slator reported on Smartling's research, noting that automated optimization narrowed the gap with expert prompts for terminology and translation, while expert judgment still led on LQA error detection.

More broadly, the conference reinforced that prompt design is not a solved problem. Structure matters more than language: a full role + context + audience + purpose prompt consistently reduced style errors across studies. Target-language prompts were preferred by expert linguists over source-language prompts. And automated optimization doesn't always win; it depends heavily on the task.

 

Low-resource languages remain the hardest frontier

Despite significant progress in high-resource language pairs, low-resource and minority languages remain a hard, largely unsolved problem. LLMs still conflate minority varieties with dominant standards, with Valencian collapsing into standard Catalan as one example that came up. Synthetic parallel data generation and RAG-based knowledge distillation are the most promising approaches, but new datasets for under-resourced languages are still scarce.

NMT isn't dead in this space either. For data-rich institutional contexts, specialized models can still win, primarily because they have access to massive, domain-specific training data that makes the investment worthwhile.

 

Evaluation is the bottleneck being discussed

Generating translations has become an easier problem to solve. Reliably defining and measuring their quality is where the field is stuck.

Traditional metrics like BLEU and TER are increasingly insufficient for real-world use cases. A wrong gender choice may barely affect TER but be a critical error in context. Automated metrics, including traditional metrics and LLM-as-a-judge approaches, still fail to reliably capture document-level quality. And as one presenter put it: without better evaluation, the industry risks optimizing in the wrong direction.

 

In-house deployment is becoming a real priority

Another major theme was the growing strategic importance of in-house and on-premise LLMs. Government, legal, parliamentary, and enterprise workflows often simply can't send sensitive content to external APIs. Data residency requirements are real, and they're not going away.

This creates a genuine market for smaller, deployable locally, specialized models, not because they're necessarily better than frontier systems in raw quality, but because they're controllable. Deployability, glossary support, style-guide handling, translation memory integration, and a credible security story are real differentiators for a large segment of the market.

 

What we're taking back with us

EAMT 2026 confirmed something we feel pretty strongly about at Smartling: the future of production translation is not just better models. It's controllable, auditable, domain-specific, and secure translation infrastructure built on top of strong LLMs, grounded with retrieval and translation memory, adapted to specific domains, and refined by humans for high-stakes content. We are also on the right track in terms of working on the open problems, including evaluation, terminology, style, and low-resource languages.

We're grateful to EAMT for creating the space for these conversations, and to the researchers and practitioners who shared their work so openly. See you next year.

 

Questions about our research or how Smartling is approaching these challenges? Póngase en contacto con nosotros.

 

¿Por qué esperar para traducir de manera más inteligente?

Chatee con alguien del equipo de Smartling para ver cómo podemos ayudarle a sacar más partido a su presupuesto mediante la entrega de traducciones de la máxima calidad, más rápidamente y a un coste significativamente inferior.
Cta-Card-Side-Image