Technical compatibility guide

Bring the real file. Keep the format.

LanguageOps extracts translation-ready content from office files, publishing packages, structured data, bilingual formats, websites, repositories and media workflows. We preserve the original structure wherever the format supports a round trip, and give you previews or filter controls when extraction needs judgement.

  • SDLXLIFF & SDLPPX
  • IDML preflight
  • Complex spreadsheets
  • Custom extraction rules

Files containing millions of words are supported. Format retention is always our priority.

Original
kept as the reconstruction base
Preview
before complex extraction
Round trip
structure, tags and metadata first
Millions
of words in one project
Supported inputs

File types we extract and process

The extension tells us what the file is. Its structure, your extraction choices and the required output determine the safest route through the platform.

Documents & publishing

DOC · DOCX · RTF · ODT · IDML · PDF

Office and publishing filters protect formatting runs, inline objects and document structure. IDML adds layer analysis and layout-aware context; scanned or visually complex PDFs can use OCR.

Example: A multilingual product manual keeps headings, emphasis, links, tables and inline formatting attached to the translated output.

Presentations

PPT · PPTX · ODP

Slide text is extracted through structure-aware office filters so translated content can be merged back into the presentation rather than delivered as a flat text list.

Example: Translate slide titles, body copy and supported notes while retaining the deck as the delivery format.

Spreadsheets & tables

XLS · XLSX · XLSM · ODS · CSV · TSV

Use a standard spreadsheet filter or the native column-pair engine for multilingual columns, repeated source/target pairs, row rules, colours, notes, context and character limits.

Example: Map C→D for product titles and E→F for descriptions, translate only blank targets, and write each result back to its original cell.

Web, software & structured content

HTML · HTM · XML · JSON · YAML · YML · PO · POT · PROPERTIES

Structure-aware filters separate translatable values from keys, markup and code. JSON key rules, HTML element previews and custom include/exclude rules help constrain extraction.

Example: Translate visible HTML copy and selected attributes while scripts, styles, element structure and protected values remain untouched.

Text, markup & subtitles

TXT · LOG · MD · MARKDOWN · RST · TEX · LATEX · SRT · VTT · SUB · SBV · ASS · SSA

Plain and marked-up text can be segmented directly or routed through a specialist filter. LaTeX commands are protected as inline codes; subtitle timing remains connected to the text workflow.

Example: Translate a LaTeX chapter while commands, references and protected arguments stay available as non-translatable tags.

Bilingual localisation files

XLIFF · XLF · XLIF · SDLXLIFF · MXLIFF · MQXLIFF

Existing bilingual files are imported directly, including source and target content, segment state and supported vendor metadata. The original bilingual is retained as the export base.

Example: Review an SDLXLIFF in the browser, move formatting tags with the translated text, and return a valid SDLXLIFF to the Trados workflow.

Packages & learning content

ZIP · SDLPPX · SDLRPX · SCORM ZIP

Package-aware inspection identifies the translatable files and supporting project metadata before processing. SCORM workflows retain course structure and connect text, subtitles and embedded media.

Example: Open an SDLPPX, detect the language direction from SDLPROJ, translate only the target-side bilingual files, and receive a rebuilt SDLPPX for delivery.

Images, audio & video

PNG · JPG · JPEG · GIF · WEBP · TIFF · BMP · MP4 · MOV · AVI · WEBM · MKV · MP3 · WAV · M4A · AAC · OGG · FLAC

OCR recovers text from images and complex PDFs. Dedicated audiovisual workflows handle transcription, subtitle translation, review, text-to-speech and dubbed delivery.

Example: Generate a timestamped transcript from video, translate it in the CAT editor, then export subtitles or a dubbed video.

Some legacy or unusually structured variants may require a filter choice or sample inspection. We confirm the extraction and delivery route rather than assuming that identical extensions contain identical structures.

Trados package compatibility

Every SDLXLIFF detail is kept, because your tools depend on it

LanguageOps recognises standalone SDLXLIFF and ZIP-backed SDLPPX or SDLRPX packages. Package inspection reads the embedded SDLPROJ to establish the real source and target variants and identify the target-side bilingual files.

SDLPPX → SDLPROJ → SDLXLIFF

  • Inline tags stay movable

    SDL <g> formatting spans and <x/> standalone codes appear as protected editor tags. They are placed with the translation and reconstructed as real XML during export.

  • Segment detail is retained

    Source and existing target, confirmation state, match percentage, origin, lock information and other supported metadata are imported at marker level.

  • Unsafe export is blocked

    A missing, substituted or unbalanced SDL tag blocks export instead of silently flattening the bilingual. The original SDLXLIFF remains the reconstruction base.

  • SDL project rules are useful

    Rules discovered in SDLPROJ can be reviewed for import into project prompts or deterministic QA. Embedded resources and remote TM or termbase references are reported separately.

Package delivery. LanguageOps retains the original SDLPPX or SDLRPX, replaces its target-side SDLXLIFF files with the completed bilinguals, and returns the rebuilt package with SDLPROJ, source bilinguals, reports and other package members preserved. Remote Trados memories or termbases referenced by SDLPROJ still require access credentials or a separately supplied TMX/TBX payload.

Complex files

Choose what is translatable before the full job runs

  • Publishing

    IDML layer and frame preflight

    Inspect layers, text-story counts, representative text and applied-language hints. Exclude hidden, alternate-language or non-deliverable layers on a working copy while the untouched IDML remains available for final merge.

  • Web

    HTML extraction preview

    Review sample text by element (headings, paragraphs, lists and table cells) before processing. Scripts and styles are excluded, while the round-trip filter protects markup.

  • Data

    Spreadsheet and CSV preview

    Verify the actual sample segments produced by selected sheets, source/target columns, column pairs, blank-target rules, row conditions and colour filters before committing.

  • Structured

    JSON key preview

    Inspect which keys and values match include/exclude rules, whether standalone strings are extracted and whether HTML-like values should be skipped.

  • Office

    Document sample preview

    Preview sample paragraphs and supported header/footer choices for Word-family files before the complete structure-aware conversion.

  • Scans

    OCR visual preview

    Review recovered content as rendered HTML before moving complex PDF or image text into translation and export workflows.

Extraction controls

Automatic when it is obvious. Configurable when it is not.

A filter determines what becomes translatable and what must remain protected. The resulting bilingual is then used by the editor, memory, terminology and QA workflow.

  1. Automatic format detection

    Known extensions select an established round-trip filter for the file type. This is the quickest route for a conventional DOCX, PPTX, HTML, PO or similar source.

    DOCX → Office filter → tagged bilingual → translated DOCX
  2. Filter options

    Choose supported format-specific behaviour such as Office comments and hidden text, PO bilingual mode, JSON key handling, OpenDocument notes or IDML extraction preferences.

    DOCX + translate comments, but leave hidden text untouched
  3. Native file engineering

    Use purpose-built extraction where ordinary filters are too limited: complex sheets, CSV column pairs, colour/row rules, IDML layer control, OCR and very large structured files.

    XLSX + three column pairs + two target languages + blank targets only
  4. Plain-language custom filters

    Describe the content you need, for example “only rows where Status is Ready” or “exclude internal_note values”. LanguageOps resolves that into a reviewable structured rule instead of running arbitrary code.

    JSON + exclude keys matching debug_* and internal_note
  5. Reusable recipes

    Save a successfully applied extraction configuration for an organisation or project so recurring files use the same reviewed rules.

    Reuse the approved supplier-catalogue mapping every month
Imports & connected sources

The source does not have to begin as a desktop file

Connected content is converted into the same controlled segment workflow, with language, context and external identity retained for review and delivery.

  • Lokalise

    Import projects and language content from Lokalise, work with translation memory, terminology and QA in LanguageOps, and keep the imported project relationship visible.

  • WordPress

    Extract titles, body content and supported fields from WordPress into bilingual segments, then publish approved translations through the connected workflow.

  • GitHub

    Select localisation files from a repository, process supported structured, text and subtitle formats, and return translated content to the development workflow.

  • Design and TMS workflows

    Figma and other localisation-system connectors can bring content into the editor with external context. XLIFF, TMX and TBX remain available for standards-based interchange.

Large-file processing

Millions of words are OK

LanguageOps is designed to keep unusually large jobs together instead of forcing you to split the source into artificial chunks.

Large JSON is streamed in bounded batches, and spreadsheet extraction retains physical row and cell locations for write-back. The editor loads work incrementally while server-side search, assignments, memory and progress operate across the complete project.

Read the large-file guide
2,000,000
word JSON example
500,000
segments in one project
250,000
row Excel example
Delivery principle

Format retention is always the priority

Where the source format supports a round trip, LanguageOps keeps the original file or bilingual skeleton and changes only the approved target content. Tags, placeholders, keys, cell locations, timing and vendor metadata are protected according to the format. If a faithful export cannot be produced, we prefer an explicit error over a deceptively successful flattened file.

  • Original document retained for reconstruction
  • Inline tags and placeholders validated
  • Native spreadsheet and CSV write-back
  • Bilingual metadata preserved where supported
  • Export errors fail visibly rather than silently flattening
Technical FAQ

Questions we are usually asked before the first upload

Does every format use the same extraction filter?
No. Conventional document and localisation formats usually use a structure-aware round-trip filter. Complex spreadsheets, OCR, IDML layer selection and very large structured sources have purpose-built extraction routes. Existing XLIFF variants are imported directly.
Can I see what will be extracted first?
Yes for the formats where configuration most affects the result: spreadsheets and CSV, Word documents, HTML, text and Markdown have extraction samples; JSON has a key/filter preview; IDML has layer and frame preflight; OCR provides a rendered HTML preview.
Can LanguageOps process an SDLPPX?
Yes. LanguageOps recognises SDLPPX and SDLRPX as packages, reads SDLPROJ metadata, identifies target-side SDLXLIFF files and imports supported segment states, tags and metadata. Completed delivery rebuilds the original package by overwriting those target bilinguals while preserving SDLPROJ, source files, reports and other untouched members.
What happens to formatting tags in SDLXLIFF?
Supported SDL formatting and standalone codes appear as protected movable tags in the source and target. Export reconstructs them as XML and blocks targets with missing, substituted or unbalanced tags.
Can you work with files containing millions of words?
Yes. LanguageOps has example projects containing two million words and 500,000 segments, and is designed to process large structured sources without manually splitting the original into smaller jobs.
Are TMX and TBX translation files?
They are linguistic-resource exchange formats. TMX imports translation-memory entries and TBX supplies terminology; they can be attached to projects but are not treated as ordinary source documents.
What if our format or extraction rule is unusual?
Use a custom filter instruction, a reusable file-engineering recipe or contact us with a representative sample. We can inspect the structure and confirm the safest extraction and reconstruction route before production work starts.

Have a difficult file?

Send us a representative sample and the required delivery format. We will confirm the extraction route, preview options and round-trip expectations before you commit the full project.