Example of PDF to Word in ASP.NET Core Smart Data Extractor Library
This sample demonstrates how PDF content can be converted into Word or HTML format while preserving structured elements such as headings, paragraphs, tables and images.
OR
Drop files (Image, PDF)
The Smart Data Extractor converts uploaded PDF documents into Word or HTML format according to the options you choose. Internally, it uses the Syncfusion OCR Processor to extract text from scanned images, with English set as the default language. For Word or HTML generation, the extractor uses the Syncfusion DocIO Library to convert raw PDF data into valid Word or HTML content. Users can download the converted files in DOCX or HTML format for further analysis. The converted Word or HTML document preserves structured elements such as headings, paragraphs, tables, and images. You can interact with the sample as follows:
- Use the Browse button to select any file of interest.
- Alternatively, drag and drop a chosen file into the designated file pick area.
- After selecting a valid file and configuring the conversion options, tap the Convert button to process the content. The converted Word or HTML file, preserving headings, paragraphs, tables and images can then be downloaded in DOCX or HTML format for further use in text-based workflows.
- Support for various file formats, including:
- PDF - '.pdf'
- Image - '.jpeg','.jpg','.png'