PDFlib TET adds performance enhancements

Version 4.0 supports Unicode postprocessing plus right-to-left and bidirectional text extraction.

August 19, 2010

Feature Release

PDFlib TET (Text Extraction Toolkit) is software for reliably extracting text information from any PDF file. It is available as a library/component and as a command-line tool. PDFlib TET makes available the text contents of a PDF as Unicode strings or structured XML, plus detailed glyph and font information. With PDFlib TET you can retrieve the corresponding Unicode values for text in a PDF document, as well as its position on the page.

Updates in V4.0

New features in PDFlib TET 4.0:

Performance enhancements: faster for many classes of documents
Higher speed and smaller memory consumption for very large documents up to hundreds of thousands of pages
Extract right-to-left and bidirectional text for Arabic, Hebrew, etc.
Unicode post-processing:
Foldings preserve, remove or replace characters
Decompositions replace a character with an equivalent sequence, e.g. replace narrow or vertical Japanese characters with their standard counterparts.
Text can be converted to all four Unicode normalization forms, e.g. emit NFC form to meet the requirements for Web text or a database.
Improved shadow removal, word boundary detection, and dehyphenation
Improved super and subscript detection
Workarounds for non-conforming PDF documents to enhance robustness
Enhanced repair mode for successfully extracting text from damaged PDF
More information in TET's XML output (TETML), e.g. dehyphenation, dropcap, shadow, and super/subscript
Improved C++ and Perl language bindings

New features in PDFlib TET PDF IFilter 4.0:

Takes advantage of the improved TET 4.0 kernel
Automatic language detection for improved search results (find word stems, partial matches, etc.)
Support for SharePoint 2010

About PDFlib

Munich-based PDFlib GmbH, founded in 2000, develops and sells leading edge components for server-centric generation and processing of PDF documents. PDFlib customers use the software for automated and high volume generation and processing of PDF documents in business and prepress workflows or for online billing systems. The components produced by PDFlib GmbH are readily available for all common environments (operating systems and programming languages) as commercial versions and partially as open source. PDFlib GmbH sells worldwide with main markets in North America, Germany and Japan.

Text extracted from a PDF with PDFlib TET.

PDFlib TET

Text extraction toolkit.

Buy Now

Got a Question?

Live Chat with our PDFlib licensing specialists now.

Official Distributor since 2003

Search Components, Applications, Add-ins and Cloud Services

Component Categories

Component Types

Component Environments

Component Publishers

1700+ Software Components in One Place

Application Categories

Application Types

Application Publishers

600+ Software Applications in One Place

Add-in Categories

Add-in Types

Add-in Publishers

250+ Software Add-ins in One Place

Bestselling Brands

200+ Publisher Brands in One Place

News by Category

News by Architecture

News by Brand

24,000+ News Articles

PDFlib TET adds performance enhancements

Updates in V4.0

About PDFlib

PDFlib TET

Got a Question?

Official Supplier

24/5 Customer Service

Trusted for 30 years

Customer Service

My Account

Company Information

Sales & Support: