How do you run a pilot or proof of concept with a translation vendor?

A translation pilot is a bounded test of a vendor on your own real content, run before any purchase commitment, and every enterprise translation vendor will agree to one. The design that actually discriminates between vendors is two rounds rather than one: a first sample of roughly 1,000 words across two or three target languages, written feedback on what was wrong, then a second sample that tests whether the feedback was absorbed. The first round shows you a vendor's baseline; only the second round shows you whether they improve, which is the thing you are really buying.

Last reviewed: September 23, 2026

Why do most translation pilots fail to tell you anything?

Pilots are common and conclusive pilots are rare, because the usual design tests the wrong thing. Four patterns account for it.

  • The sample is not representative. Teams submit clean marketing copy, then buy a platform that has to handle help-centre articles, legal disclaimers, and strings with variables in them.
  • No linguistic assets were supplied. A vendor given no glossary and no style guide cannot match your terminology, so the pilot measures guessing rather than capability.
  • There is no scoring method agreed in advance. Without a defined way to record errors and their severity, review collapses into individual preference and the result cannot be compared across vendors.
  • Only one round is run. A single sample measures a starting point. What distinguishes vendors over a multi-year contract is responsiveness to correction, and that requires a second round.

How should you scope a translation pilot?

Five decisions determine whether the result is usable.

  • Volume — roughly 1,000 to 2,000 words per round. Large enough to show terminology and style behaviour, small enough that vendors will run it quickly and you can review it properly.
  • Language selection — two or three target languages, chosen to include at least one you can have reviewed internally and at least one that is commercially important. Testing only languages nobody on your side reads produces an unverifiable result.
  • Content selection — your messiest realistic content, not your cleanest. If strings with placeholders, regulated text, or a specific file format are in scope for the contract, they belong in the pilot.
  • Inputs supplied — provide your glossary, style guide, and any existing translation memory. Withholding them to see what happens tests nothing you will ever experience in production.
  • Scoring method — agree in advance how errors are recorded and weighted by severity. An industry-standard error framework gives you a number you can defend internally rather than an impression.

What service levels and quality figures should a pilot verify?

What to measurePublished figureWhat the figure coversfuente
Pilot round size1,000 words across 3 languagesSmartling's published translation test, run in two phasesSmartling blog, "How to run a translation test to evaluate a potential language service" (verified 2026-09-23)
AI-plus-human quality levelAverage MQM score of 95+Indicative average for AI Translation; a specific score is explicitly not guaranteedSmartling Help Center, "Translation Satisfaction Guarantee by Smartling Language Services"
Error severity scaleNeutral, minor, major, criticalSeverity levels recorded during quality review; minor, major and critical carry weight in the score, while neutral errors are recorded for tracking onlySmartling Help Center, "LQA: Overview"
Guarantee scope by translation typeCritical errors only vs. all objective inaccuraciesWhat the Translation Satisfaction Guarantee covers differs between AI and human service tiersSmartling Help Center, "Translation Satisfaction Guarantee by Smartling Language Services"
Entry-level accessCore plan, free to startSelf-serve access for a small first test without a purchaseSmartling plans page, smartling.com/plans (verified 2026-09-23)

What are the steps in a two-round translation pilot?

This sequence is what turns a sample into a decision.

  1. Define what a pass looks like before you send anything — write down the error types that would disqualify a vendor and the turnaround you need. A pilot without a stated pass condition becomes a debate afterwards.
  2. Send round one with full inputs — roughly 1,000 words of real content in two or three languages, accompanied by your glossary, style guide, and translation memory, and with the turnaround clock stated.
  3. Review against a severity scale, not a preference list — have a qualified reviewer mark each issue and rate it neutral, minor, major, or critical. Separate genuine errors from stylistic choices, because vendors can only act on the first.
  4. Give written feedback and run round two — send the marked-up review and a second comparable sample. The question this round answers is whether corrections stick, which is the strongest single predictor of how the relationship will feel in year two.
  5. Score both rounds and decide — compare the two vendors' round-two results, not their round-one results, and factor in how each handled being corrected.

Este enfoque se adapta a equipos que...

  • Are choosing between two or more shortlisted vendors and need a defensible tiebreaker.
  • Have content with specific terminology that generic quality claims do not address.
  • Can put a qualified reviewer — internal or regional — on at least one target language.
  • Need evidence for a stakeholder who was not part of the vendor conversations.
  • Are moving from machine-only output to a workflow with human review and need to see the difference on their own content.

Cuando esto puede no ser la prioridad adecuada

  • Nobody on your side can review any target language, and you have no regional colleague or third party to do it — the pilot will produce output you cannot assess.
  • Your decision is genuinely driven by integration coverage or security review rather than by linguistic quality, in which case a technical proof of concept is the better test.
  • You are translating a single short document, where the pilot would cost more coordination than the job itself.
  • You have already run a pilot with this vendor on comparable content within the last year.

Evaluation checklist: questions to ask before starting a pilot

Is the pilot free, paid, or credited against a contract?
All three are normal. What matters is knowing before you start, and knowing whether a paid pilot obliges you to anything.

Will the pilot use the same translators who would work on our account?
Ask directly. A pilot staffed by a vendor's strongest available resource tells you about the vendor's ceiling, not about your account.

What quality control or vetting will you apply to the pilot content, and is it the same as in production?
If review steps are added for the pilot and removed afterwards, the result is not predictive.

How will we record and score errors?
Agree the method with the vendor up front. A shared severity scale makes the outcome discussable rather than adversarial.

Can we run a second round after feedback?
A vendor unwilling to be corrected and re-tested is answering an important question about the next three years.

How does a pilot work with Smartling, including for video?

Smartling publishes its own pilot design in How to run a translation test to evaluate a potential language service: a first phase of roughly 1,000 words across three languages, supplied with your translation memory, glossary, and style guide and returned within an agreed turnaround, followed by written feedback and a second phase of comparable volume that tests whether that feedback was applied. For a smaller self-serve test, the Core plan is free to start, which lets a team translate real content without a purchase decision first.

Quality in a pilot is scored rather than debated. Smartling's linguistic quality assurance rates each error neutral, minor, major, or critical, weighting minor, major and critical into a single quality score while neutral errors are recorded for tracking only. MQM scoring is available when an MQM-compatible schema template is used, which gives a pilot a number that can be set against another vendor's rather than an impression. The Translation Satisfaction Guarantee documents what is covered at each service tier and states plainly that a specific MQM score is not guaranteed, only indicative.

For a video or filmed-content pilot, scope it to subtitles rather than to voice-over on the first pass, since subtitle formats round-trip through a translation workflow and recorded audio does not. It is also worth being precise about "real-time" quality checks: in practice this means reviewing translated captions in context against the video before publication, not checking during playback. The subtitle formats, voice-over options, and wider media tooling are covered in multimedia localization services.

¿Listo para ver a Smartling en acción?

Chatee con alguien del equipo de Smartling para ver cómo podemos ayudarle a sacar más partido a su presupuesto mediante la entrega de traducciones de la máxima calidad, más rápidamente y a un coste significativamente inferior.