Browse our Products
If so you can download any of the below versions for testing. The product will function as normal except for an evaluation limitation. At the time of purchase we provide a license file via email that will allow the product to work in its full capacity. If you would also like an evaluation license to test without any restrictions for 30 days, please follow the directions provided here.
If you experience errors, when you try to download a file, make sure your network policies (enforced by your company or ISP) allow downloading ZIP and/or MSI files.
Powerful document parsing API for Node.js (powered by Java) to extract text, metadata, images, tables, barcodes, hyperlinks and structured data from 50+ document formats, including Microsoft Office and OpenDocument types, PDF documents, emails, e-books and archives. Parse documents by user-defined templates, read PDF forms, generate page previews and export extracted data to XML or JSON.
See the full list of supported formats.
To install the package, check the System Requirements and Installation documentation topics for platform-specific instructions.
npm install @groupdocs/groupdocs.parser
Here are some typical use cases:
This example shows how to extract a text from a DOCX file.
'use strict';
const groupdocs = require('@groupdocs/groupdocs.parser');
// Apply license, required for non-evaluation usage
const license = new groupdocs.License();
license.setLicense("GroupDocs.Parser.lic");
// Extract text
const parser = new groupdocs.Parser("document.docx");
const reader = parser.getText();
if (reader === null) {
console.log("Text extraction isn't supported");
} else {
console.log(reader.readToEnd());
reader.close();
}
parser.close();
// Exit
process.exit(0);
This example prints the metadata of a DOCX file.
'use strict';
const groupdocs = require('@groupdocs/groupdocs.parser');
const parser = new groupdocs.Parser("document.docx");
// Iterate over metadata items
const iterator = parser.getMetadata().iterator();
while (iterator.hasNext()) {
const item = iterator.next();
console.log(`${item.getName()}: ${item.getValue()}`);
}
parser.close();
process.exit(0);
This example reads the values of all form fields of a PDF document.
'use strict';
const java = require('java');
const groupdocs = require('@groupdocs/groupdocs.parser');
const parser = new groupdocs.Parser("form.pdf");
// Extract form data
const data = parser.parseForm();
for (let i = 0; i < data.getCount(); i++) {
const field = data.get(i);
const area = field.getPageArea();
const isText = area !== null && java.instanceOf(area, 'com.groupdocs.parser.data.PageTextArea');
console.log(`${field.getName()}: ${isText ? area.getText() : "Not a text field"}`);
}
parser.close();
process.exit(0);
GroupDocs.Parser GroupDocs.Total Conholdate Conholdate.Total Node.js NPM Document-Parsing Text-Extraction Metadata Images Barcodes Tables Hyperlinks PDF DOC DOCX XLS XLSX PPT PPTX EML MSG ZIP eBook