Product updates · · Frank Wei
XLIFF 1.2: the CAT interchange format BabelBee now speaks
Why XLIFF is the safest bridge between professional CAT tools and an AI-native translation production platform — and how to run a round-trip without losing inline tags.
If your LSP or language team already uses a computer-assisted translation (CAT) tool, you probably do not want to abandon segment boundaries, inline tags, or the QA habits that come with them. You also may not want every PM or reviewer to buy a CAT seat just to run one AI-assisted production pass.
XLIFF 1.2 is the common answer. It is an open, vendor-neutral exchange format that Trados, memoQ, Phrase, OmegaT, and many other tools can export and re-import. BabelBee now supports .xliff and .xlf end to end — so you can keep CAT as the system of record while using an AI-native production workbench for translation and review.
What XLIFF actually carries
An XLIFF file is XML. Each <trans-unit> is one translatable segment. The <source> element holds the source text plus inline markup — formatting tags, placeholders, links, and other non-translatable structure that CAT tools represent as <g>, <x>, <bx>, and similar elements.
That is the key difference from uploading a plain Word file after CAT has injected tags into the document body. XLIFF keeps structure explicit and machine-readable.
Why not round-trip a tagged Word file?
Many CAT workflows produce bilingual or tagged DOCX files. Those files look fine in Word, but they embed vendor-specific styles and inline tag representations. A generic document pipeline that reads "visible text" and writes translations back into runs will flatten or break those tags.
BabelBee detects common CAT-tagged DOCX patterns and:
- Warns at segmentation review — so you know format-restore export is risky
- Blocks format-restore export on those files — rather than silently corrupting tags
The supported path for tagged content is XLIFF in, XLIFF out.
How the BabelBee XLIFF workflow works
- Export XLIFF 1.2 from your CAT tool (one file per language pair is typical).
- Upload to BabelBee — accepted on the same upload page as DOCX, PDF, and other formats.
- Automatic segmentation — each
<trans-unit>becomes one segment. Inline tags are converted to stable placeholders such as[[g:1]]…[[/g:1]]or[[x:2/]]so the AI can translate around them. - Translate and review — AI translation, second-model QA, and optional human review run in the workbench like any other project.
- Export XLIFF — translated
<target>elements are written back with inline tags restored. - Re-import into your CAT tool — run your usual TM update, LQA, and delivery steps.
You get AI speed in the middle without giving up CAT-native structure at the edges.
What we preserve today
BabelBee's XLIFF 1.2 support focuses on the inline elements teams hit most often in real exports:
- Paired tags such as
<g>(generic group / formatting) - Standalone placeholders such as
<x/> - Common CAT inline types including
bx/ex,bpt/ept,ph,it,mrk
During translation, placeholders must survive unchanged. Automated QA flags missing or altered [[…]] tokens, and markup-heavy batches route through the LLM engine that can follow placeholder instructions reliably.
A realistic example
Source segment in XLIFF:
<trans-unit id="42">
<source>Click <g id="1">here</g> to continue.</source>
</trans-unit>
Inside BabelBee the segment appears as:
Click [[g:1]]here[[/g:1]] to continue.
After translation to German, export might produce:
<trans-unit id="42">
<source>Click <g id="1">here</g> to continue.</source>
<target state="translated">Klicken Sie <g id="1">hier</g>, um fortzufahren.</target>
</trans-unit>
Your CAT tool sees familiar structure — not a flattened sentence with broken tags.
When XLIFF is the right choice
Use XLIFF when:
- Documents contain inline tags or variables that must survive translation
- You already manage TM and terminology in a CAT suite
- Reviewers expect to finish QA in CAT after an AI draft
- You need a documented handoff between tools, not a one-off copy-paste
Skip XLIFF when:
- The file is plain marketing copy or internal notes with no markup
- You only need a formatted DOCX/PDF deliverable and never return to CAT
For those simpler cases, uploading DOCX or PDF directly is still the fastest path.
Frequently asked questions
Which XLIFF version?
XLIFF 1.2 — the dialect virtually every CAT exporter still ships. XLIFF 2.x is not supported yet.
Can I split or merge segments?
You can merge awkward fragments in the segmentation review step, but the default XLIFF profile keeps one segment per trans-unit so tag boundaries stay aligned with your CAT project. Heavy re-segmentation is better done before export.
Does this replace TMX?
No. TMX carries translation memory; XLIFF carries the live document and its tags. Many teams use both — TMX for recall, XLIFF for the job at hand. See our TMX migration guide for memory workflows.
What if I only have a tagged DOCX from a client?
Ask for an XLIFF export from whoever prepared the file in CAT. If that is not possible, BabelBee can still translate the visible text, but format-restore round-trip is not supported for CAT-tagged DOCX.
Ready to try it on a real export? Start a free workspace or read about CAT vs pure AI workflows.