Documents & publishing
DOC · DOCX · RTF · ODT · IDML · PDF
Office and publishing filters protect formatting runs, inline objects and document structure. IDML adds layer analysis and layout-aware context; scanned or visually complex PDFs can use OCR.
Technical compatibility guide
LanguageOps extracts translation-ready content from office files, publishing packages, structured data, bilingual formats, websites, repositories and media workflows. We preserve the original structure wherever the format supports a round trip—and give you previews or filter controls when extraction needs judgement.
Files containing millions of words are supported. Format retention is always our priority.
Supported inputs
The extension tells us what the file is. Its structure, your extraction choices and the required output determine the safest route through the platform.
DOC · DOCX · RTF · ODT · IDML · PDF
Office and publishing filters protect formatting runs, inline objects and document structure. IDML adds layer analysis and layout-aware context; scanned or visually complex PDFs can use OCR.
PPT · PPTX · ODP
Slide text is extracted through structure-aware office filters so translated content can be merged back into the presentation rather than delivered as a flat text list.
XLS · XLSX · XLSM · ODS · CSV · TSV
Use a standard spreadsheet filter or the native column-pair engine for multilingual columns, repeated source/target pairs, row rules, colours, notes, context and character limits.
HTML · HTM · XML · JSON · YAML · YML · PO · POT · PROPERTIES
Structure-aware filters separate translatable values from keys, markup and code. JSON key rules, HTML element previews and custom include/exclude rules help constrain extraction.
TXT · LOG · MD · MARKDOWN · RST · TEX · LATEX · SRT · VTT · SUB · SBV · ASS · SSA
Plain and marked-up text can be segmented directly or routed through a specialist filter. LaTeX commands are protected as inline codes; subtitle timing remains connected to the text workflow.
XLIFF · XLF · XLIF · SDLXLIFF · MXLIFF · MQXLIFF
Existing bilingual files are imported directly, including source and target content, segment state and supported vendor metadata. The original bilingual is retained as the export base.
ZIP · SDLPPX · SDLRPX · SCORM ZIP
Package-aware inspection identifies the translatable files and supporting project metadata before processing. SCORM workflows retain course structure and connect text, subtitles and embedded media.
PNG · JPG · JPEG · GIF · WEBP · TIFF · BMP · MP4 · MOV · AVI · WEBM · MKV · MP3 · WAV · M4A · AAC · OGG · FLAC
OCR recovers text from images and complex PDFs. Dedicated audiovisual workflows handle transcription, subtitle translation, review, text-to-speech and dubbed delivery.
Some legacy or unusually structured variants may require a filter choice or sample inspection. We confirm the extraction and delivery route rather than assuming that identical extensions contain identical structures.
Trados package compatibility
LanguageOps recognises standalone SDLXLIFF and ZIP-backed SDLPPX or SDLRPX packages. Package inspection reads the embedded SDLPROJ to establish the real source and target variants and identify the target-side bilingual files.
SDL <g> formatting spans and <x/> standalone codes appear as protected editor tags. They are placed with the translation and reconstructed as real XML during export.
Source and existing target, confirmation state, match percentage, origin, lock information and other supported metadata are imported at marker level.
A missing, substituted or unbalanced SDL tag blocks export instead of silently flattening the bilingual. The original SDLXLIFF remains the reconstruction base.
Rules discovered in SDLPROJ can be reviewed for import into project prompts or deterministic QA. Embedded resources and remote TM or termbase references are reported separately.
LanguageOps retains the original SDLPPX or SDLRPX, replaces its target-side SDLXLIFF files with the completed bilinguals, and returns the rebuilt package with SDLPROJ, source bilinguals, reports and other package members preserved. Remote Trados memories or termbases referenced by SDLPROJ still require access credentials or a separately supplied TMX/TBX payload.
Complex files
Inspect layers, text-story counts, representative text and applied-language hints. Exclude hidden, alternate-language or non-deliverable layers on a working copy while the untouched IDML remains available for final merge.
Review sample text by element—headings, paragraphs, lists and table cells—before processing. Scripts and styles are excluded, while the round-trip filter protects markup.
Verify the actual sample segments produced by selected sheets, source/target columns, column pairs, blank-target rules, row conditions and colour filters before committing.
Inspect which keys and values match include/exclude rules, whether standalone strings are extracted and whether HTML-like values should be skipped.
Preview sample paragraphs and supported header/footer choices for Word-family files before the complete structure-aware conversion.
Review recovered content as rendered HTML before moving complex PDF or image text into translation and export workflows.
Extraction controls
A filter determines what becomes translatable and what must remain protected. The resulting bilingual is then used by the editor, memory, terminology and QA workflow.
Known extensions select an established round-trip filter for the file type. This is the quickest route for a conventional DOCX, PPTX, HTML, PO or similar source.
DOCX → Office filter → tagged bilingual → translated DOCXChoose supported format-specific behaviour such as Office comments and hidden text, PO bilingual mode, JSON key handling, OpenDocument notes or IDML extraction preferences.
DOCX + translate comments, but leave hidden text untouchedUse purpose-built extraction where ordinary filters are too limited: complex sheets, CSV column pairs, colour/row rules, IDML layer control, OCR and very large structured files.
XLSX + three column pairs + two target languages + blank targets onlyDescribe the content you need—for example, “only rows where Status is Ready” or “exclude internal_note values”. LanguageOps resolves that into a reviewable structured rule instead of running arbitrary code.
JSON + exclude keys matching debug_* and internal_noteSave a successfully applied extraction configuration for an organisation or project so recurring files use the same reviewed rules.
Reuse the approved supplier-catalogue mapping every monthImports & connected sources
Connected content is converted into the same controlled segment workflow, with language, context and external identity retained for review and delivery.
Import projects and language content from Lokalise, work with translation memory, terminology and QA in LanguageOps, and keep the imported project relationship visible.
Extract titles, body content and supported fields from WordPress into bilingual segments, then publish approved translations through the connected workflow.
Select localisation files from a repository, process supported structured, text and subtitle formats, and return translated content to the development workflow.
Figma and other localisation-system connectors can bring content into the editor with external context. XLIFF, TMX and TBX remain available for standards-based interchange.
Large-file processing
LanguageOps is designed to keep unusually large jobs together instead of forcing you to split the source into artificial chunks.
Large JSON is streamed in bounded batches, and spreadsheet extraction retains physical row and cell locations for write-back. The editor loads work incrementally while server-side search, assignments, memory and progress operate across the complete project.
Read the large-file guideDelivery principle
Where the source format supports a round trip, LanguageOps keeps the original file or bilingual skeleton and changes only the approved target content. Tags, placeholders, keys, cell locations, timing and vendor metadata are protected according to the format. If a faithful export cannot be produced, we prefer an explicit error over a deceptively successful flattened file.
Technical FAQ
No. Conventional document and localisation formats usually use a structure-aware round-trip filter. Complex spreadsheets, OCR, IDML layer selection and very large structured sources have purpose-built extraction routes. Existing XLIFF variants are imported directly.
Yes for the formats where configuration most affects the result: spreadsheets and CSV, Word documents, HTML, text and Markdown have extraction samples; JSON has a key/filter preview; IDML has layer and frame preflight; OCR provides a rendered HTML preview.
Yes. LanguageOps recognises SDLPPX and SDLRPX as packages, reads SDLPROJ metadata, identifies target-side SDLXLIFF files and imports supported segment states, tags and metadata. Completed delivery rebuilds the original package by overwriting those target bilinguals while preserving SDLPROJ, source files, reports and other untouched members.
Supported SDL formatting and standalone codes appear as protected movable tags in the source and target. Export reconstructs them as XML and blocks targets with missing, substituted or unbalanced tags.
Yes. LanguageOps has example projects containing two million words and 500,000 segments, and is designed to process large structured sources without manually splitting the original into smaller jobs.
They are linguistic-resource exchange formats. TMX imports translation-memory entries and TBX supplies terminology; they can be attached to projects but are not treated as ordinary source documents.
Use a custom filter instruction, a reusable file-engineering recipe or contact us with a representative sample. We can inspect the structure and confirm the safest extraction and reconstruction route before production work starts.
Send us a representative sample and the required delivery format. We will confirm the extraction route, preview options and round-trip expectations before you commit the full project.