Latest release (September 2026)
The latest release v26.9 is the first release of GroupDocs.Parser for Node.js via Java. It brings the full GroupDocs.Parser for Java 26.9 feature set to Node.js applications: extract text, metadata, images, tables, barcodes and hyperlinks from 50+ document formats, parse documents by templates and read PDF forms using JavaScript.
Full list of changes in this release
| Key | Category | Summary |
|---|---|---|
| PARSERNODEJS | Feature | Implement GroupDocs.Parser for Node.js via Java wrapper based on GroupDocs.Parser for Java 26.9 |
| PARSERNODEJS | Feature | Publish the @groupdocs/groupdocs.parser package to npm |
| PARSERNODEJS | Feature | Add helpers to load documents and licenses from Node.js streams |
| PARSERNODEJS | Feature | Add code examples for GroupDocs.Parser for Node.js via Java |
Major features
Initial product release with the following features:
- Text extraction: plain, raw and formatted (HTML, Markdown) text from the whole document or a single page
- Metadata extraction: document properties of Office documents, PDF, emails and e-books
- Images extraction: images from documents, pages and page areas with the ability to save them to files
- Tables extraction: tables from documents and pages, including tables described by a template layout
- Barcodes extraction: barcodes from documents and pages with export to XML or JSON
- Hyperlinks extraction: hyperlinks from documents, pages and page areas
- Search: search text by keyword or regular expression
- Parsing by templates: extract data using fixed, regex and linked template fields and template tables
- Forms: read field values of PDF forms
- Containers: iterate through ZIP archives, PDF portfolios, email attachments and Outlook storages
- Table of contents: extract the table of contents and the text of its items
- Page previews: generate PNG, JPEG or BMP previews of document pages
- Document information: get the file type, page count and size, and check the supported features for a document
Supported file formats:
- Microsoft Word®: DOC, DOCX, DOCM, DOT, DOTX, DOTM, RTF
- Microsoft Excel®: XLS, XLSX, XLSM, XLSB, XLT, XLTX, XLTM, XLA, XLAM, CSV
- Microsoft PowerPoint®: PPT, PPTX, PPTM, PPS, PPSX, PPSM, POT, POTX, POTM
- Microsoft OneNote®: ONE
- Microsoft Outlook®: MSG, PST, OST
- OpenDocument: ODT, OTT, ODS, OTS, ODP, OTP
- Fixed Layout: PDF
- Email: EML, EMLX
- eBook: EPUB, FB2, MOBI, AZW3, CHM
- Web: HTML, XHTML, MHTML, XML, Markdown
- Archives: ZIP, RAR, 7Z, TAR, GZ, BZ2
- Text: TXT
Package
The package is available on npm: @groupdocs/groupdocs.parser.
npm install @groupdocs/groupdocs.parser
The package requires Node.js and Java Runtime Environment (JRE) 8 or later. JRE 17 or later is required for APIs that take callbacks implemented in JavaScript (for example, page previews).
Code examples
Examples: GroupDocs.Parser for Node.js via Java — Code Examples
Public API and backward incompatible changes
The public API matches GroupDocs.Parser for Java 26.9. All public classes are exported from the package root, for example Parser, License, Metered, LoadOptions, TextOptions, SearchOptions, Template and FileType.
The following Node.js specific helpers were added:
readDataFromStream(readStream, callback)- reads a Node.js stream into ajava.io.InputStreamthat can be passed to theParserconstructorreadBytesFromStream(readStream, callback)- reads a Node.js stream into a Javabyte[]arrayStreamBuffer- accumulates Node.js stream chunks and exposes them as a Javabyte[]array orjava.io.InputStreamLicense.setLicenseFromStream(license, licenseStream, callback)- sets a license from a Node.js stream
Feedback
We value your feedback! If you have any questions, issues, or suggestions, feel free to reach out to us through our Free Support Forum. Our team will be happy to assist you and answer any questions you may have.