Ratings and Reviews 0 Ratings

Total
ease
features
design
support

This software has no reviews. Be the first to write a review.

Write a Review

Ratings and Reviews 0 Ratings

Total
ease
features
design
support

This software has no reviews. Be the first to write a review.

Write a Review

Alternatives to Consider

  • MobiOffice Reviews & Ratings
    14,822 Ratings
    Company Website
  • Foxit Document Workflow APIs Reviews & Ratings
    7 Ratings
    Company Website
  • PDF Guru Reviews & Ratings
    83,791 Ratings
    Company Website
  • MobiPDF Reviews & Ratings
    7,001 Ratings
    Company Website
  • Docmosis Reviews & Ratings
    51 Ratings
    Company Website
  • PDFCreator Reviews & Ratings
    557 Ratings
    Company Website
  • Nutrient SDK Reviews & Ratings
    111 Ratings
    Company Website
  • Concord Reviews & Ratings
    237 Ratings
    Company Website
  • Apryse PDF SDK Reviews & Ratings
    157 Ratings
    Company Website
  • Titan Reviews & Ratings
    376 Ratings
    Company Website

What is pdf2docx?

pdf2docx is a Python library that utilizes PyMuPDF to extract data from PDF files, analyze their layouts according to defined rules, and generate .docx documents using python-docx. This library simplifies the conversion of numerous elements such as text, images, and tables, featuring capabilities for table extraction, formatting management, and preservation of layout integrity whenever feasible. Additionally, it provides both a command-line interface and a graphical user interface to suit various user needs. Its modular design includes separate packages for handling pages, layouts, tables, images, shape paths, text spans, and other components, offering precise control over the transformation of PDF content into Word files. Developers can utilize the API for batch processing or easily embed it within their existing systems. Extensive documentation is available, detailing installation (which can be sourced from PyPI or directly), usage guidelines, and in-depth technical information on layout parsing, table extraction, and the internal modules. The project is open-source and can be found on GitHub, published under its license and with a disclaimer of any warranties. Furthermore, pdf2docx not only streamlines the conversion process significantly but also serves as an invaluable resource for professionals regularly working with PDF and Word file formats, enhancing their productivity.

What is Unsiloed?

Unsiloed AI is a document layer for enterprise AI that converts complex unstructured files into clean JSON, Markdown, and structured data. The platform is built for organizations whose most valuable information lives inside PDFs, scanned documents, images, spreadsheets, contracts, invoices, reports, filings, forms, and other hard-to-parse formats. Unsiloed helps AI teams avoid building brittle OCR, parser, and post-processing pipelines by providing a production-ready API for document parsing, field extraction, and document splitting. Its parsing capability converts PDFs, scans, and images into LLM-ready Markdown while preserving tables, figures, text hierarchy, page structure, signatures, handwriting, and visual context. Its extraction capability pulls specific fields into JSON using schemas, confidence thresholds, and domain-aware logic that can understand context such as line items, payment terms, clauses, totals, and references. Its splitting capability separates multi-document files into individual documents and breaks long files into retrievable chunks for RAG, agent workflows, and search systems. Unsiloed uses proprietary dual-stream vision models that process content and layout in parallel, then fuse them through cross-attention so the system can reason over what a document says and how it is structured. The platform’s architecture includes attention-guided heatmaps, typed document regions, layout-aware processing, and domain-specific decoding for industries such as finance, legal, healthcare, and enterprise operations. It is designed to handle edge cases that traditional OCR often misses, including nested tables, merged cells, multi-page tables, figures, handwritten notes, forms, and documents with mixed formats. Teams can connect data sources such as S3, SharePoint, Drive, Snowflake, or a document management system, then send structured outputs into LLMs, AI agents, vector databases, or analytics warehouses.

Media

Media

No images available

Integrations Supported

GitHub
Microsoft Word
PyMuPDF
PyPI
Python

Integrations Supported

GitHub
Microsoft Word
PyMuPDF
PyPI
Python

API Availability

Has API

API Availability

Has API

Pricing Information

Free
Free Version
Free Trial Offered?

Pricing Information

Pricing not provided
Free Version
Free Trial Offered?

Supported Platforms

SaaS
Android
iPhone
iPad
Windows
Mac
On-Prem
Chromebook
Linux

Supported Platforms

SaaS
Android
iPhone
iPad
Windows
Mac
On-Prem
Chromebook
Linux

Customer Service / Support

Standard Support
24 Hour Support
Web-Based Support

Customer Service / Support

Standard Support
24 Hour Support
Web-Based Support

Training Options

Documentation Hub
Webinars
Online Training
On-Site Training

Training Options

Documentation Hub
Webinars
Online Training
On-Site Training

Company Facts

Organization Name

Artifex

Date Founded

1993

Company Location

United States

Company Website

pdf2docx.readthedocs.io/en/latest/

Company Facts

Organization Name

Unsiloed.ai

Date Founded

2025

Company Location

United States

Company Website

www.unsiloed.ai/

Categories and Features

PDF

Annotations
Convert to PDF
Digital Signature
Encryption
Merge / Append
PDF Reader
Watermarking

Categories and Features

Data Extraction

Disparate Data Collection
Document Extraction
Email Address Extraction
IP Address Extraction
Image Extraction
Phone Number Extraction
Pricing Extraction
Web Data Extraction

Popular Alternatives

AnyParser Reviews & Ratings

AnyParser

CambioML

Popular Alternatives

PDF.co  Reviews & Ratings

PDF.co

ByteScout
PDF Conversa Reviews & Ratings

PDF Conversa

ASCOMP Software