Browse our Products

If so you can download any of the below versions for testing. The product will function as normal except for an evaluation limitation. At the time of purchase we provide a license file via email that will allow the product to work in its full capacity. If you would also like an evaluation license to test without any restrictions for 30 days, please follow the directions provided here.

GroupDocs.Parser for Node.js via Java Downloads

Powerful document parsing API for Node.js (powered by Java) to extract text, metadata, images, tables, barcodes, hyperlinks and structured data from 50+ document formats, including Microsoft Office and OpenDocument types, PDF documents, emails, e-books and archives. Parse documents by user-defined templates, read PDF forms, generate page previews and export extracted data to XML or JSON.

Key Features

  • Extract plain, raw or formatted (HTML, Markdown) text from the whole document or a single page.
  • Extract metadata, images, tables, hyperlinks, barcodes and table of contents.
  • Search text by keyword or regular expression.
  • Parse documents by templates and read PDF form fields.
  • Iterate through ZIP archives, PDF portfolios and email attachments.
  • Generate document page previews and export data to XML or JSON.

Supported Formats

Microsoft Office Formats

  • Microsoft Word: DOC, DOCX, DOCM, DOT, DOTX, DOTM, RTF
  • Microsoft Excel: XLS, XLSX, XLSM, XLSB, XLT, XLTX, XLTM, XLA, XLAM, CSV
  • Microsoft PowerPoint: PPT, PPTX, PPTM, PPS, PPSX, PPSM, POT, POTX, POTM
  • Microsoft OneNote: ONE
  • Microsoft Outlook: MSG, PST, OST

Other Supported Formats

  • OpenDocument: ODT, OTT, ODS, OTS, ODP, OTP
  • Portable: PDF
  • Email: EML, EMLX
  • eBook: EPUB, FB2, MOBI, AZW3, CHM
  • Web: HTML, XHTML, MHTML, XML, Markdown
  • Archives: ZIP, RAR, 7Z, TAR, GZ, BZ2
  • Text: TXT

See the full list of supported formats.

Getting Started

Prerequisites

  • Node.js (LTS recommended)
  • Java Runtime Environment (JRE) 8 or later; JRE 17 or later is required for APIs that take callbacks implemented in JavaScript (e.g. page previews)
  • Windows, Linux, or macOS

Installation

To install the package, check the System Requirements and Installation documentation topics for platform-specific instructions.

npm install @groupdocs/groupdocs.parser

Use cases

Here are some typical use cases:

Extract text

This example shows how to extract a text from a DOCX file.

'use strict';

const groupdocs = require('@groupdocs/groupdocs.parser');

// Apply license, required for non-evaluation usage
const license = new groupdocs.License();
license.setLicense("GroupDocs.Parser.lic");

// Extract text
const parser = new groupdocs.Parser("document.docx");
const reader = parser.getText();
if (reader === null) {
  console.log("Text extraction isn't supported");
} else {
  console.log(reader.readToEnd());
  reader.close();
}
parser.close();

// Exit
process.exit(0);

Extract metadata

This example prints the metadata of a DOCX file.

'use strict';

const groupdocs = require('@groupdocs/groupdocs.parser');

const parser = new groupdocs.Parser("document.docx");

// Iterate over metadata items
const iterator = parser.getMetadata().iterator();
while (iterator.hasNext()) {
  const item = iterator.next();
  console.log(`${item.getName()}: ${item.getValue()}`);
}

parser.close();
process.exit(0);

Parse PDF form fields

This example reads the values of all form fields of a PDF document.

'use strict';

const java = require('java');

const groupdocs = require('@groupdocs/groupdocs.parser');

const parser = new groupdocs.Parser("form.pdf");

// Extract form data
const data = parser.parseForm();
for (let i = 0; i < data.getCount(); i++) {
  const field = data.get(i);
  const area = field.getPageArea();
  const isText = area !== null && java.instanceOf(area, 'com.groupdocs.parser.data.PageTextArea');
  console.log(`${field.getName()}: ${isText ? area.getText() : "Not a text field"}`);
}

parser.close();
process.exit(0);


Direct Download