shipfastlabs/
parsel
g4t.io

Parsel: Parse PDFs, Office Documents, and Images in PHP

363 stars by Yannick Lyn Fatt

An expressive, driver-based PHP API for local and remote document parsers.

README from shipfastlabs/parsel Open on GitHub →

Parsel

Introduction

Parsel provides an expressive PHP API for extracting Markdown, text, and structured data from documents. It ships with drivers for the local LiteParse and AnyDoc command-line parsers, and you may register your own.

use Shipfastlabs\Parsel;

$markdown = Parsel::file('report.pdf')->markdown();

Installation

Parsel requires PHP 8.4 or greater. You may install it via Composer:

composer require shipfastlabs/parsel

Next, install the parser binary for the driver you plan to use. By default, the installer installs LiteParse:

vendor/bin/parsel-install
vendor/bin/parsel-install --driver=anydoc
vendor/bin/parsel-install --driver=all

You may choose a package manager with --manager. LiteParse supports npm, pnpm, bun, pip, and cargo; AnyDoc supports npm, pnpm, and bun and requires Node.js 20 or greater. LiteParse needs LibreOffice to convert Office, OpenDocument, RTF, and CSV files, which --with-system-dependencies installs along with ImageMagick:

vendor/bin/parsel-install --driver=anydoc --manager=pnpm
vendor/bin/parsel-install --with-system-dependencies

Basic Usage

Sources

You may parse a document from a path or from raw bytes:

use Shipfastlabs\Parsel;

$markdown = Parsel::file('/path/to/report.pdf')->markdown();

$markdown = Parsel::bytes($contents, 'pdf')->markdown();

[!NOTE] Byte sources always require a file extension so the parser can identify the format.

Output

Every driver can return Markdown. LiteParse can also return plain text, a structured Document, or an array:

$markdown = Parsel::file('report.pdf')->markdown();
$text = Parsel::file('report.pdf')->text();
$array = Parsel::file('report.pdf')->toArray();

$document = Parsel::file('report.pdf')->parse();

echo $document->pageCount();

foreach ($document->pages as $page) {
    foreach ($page->items as $item) {
        echo "{$item->text} @ ({$item->x}, {$item->y})";
    }
}

For large documents, the lazyPages method streams pages one at a time:

foreach (Parsel::file('large.pdf')->lazyPages() as $page) {
    echo $page->text;
}

The screenshots method renders each page to a PNG in an existing directory and returns the file paths:

$paths = Parsel::file('report.pdf')->screenshots('/path/to/screenshots');

The save method writes the output to disk, choosing the format from the extension: .md and .markdown save Markdown, .json saves structured JSON, and any other extension saves plain text:

Parsel::file('report.pdf')->save('report.md');
Parsel::file('report.pdf')->save('report.json');

Timeouts

Parsing times out after 60 seconds by default. You may change the timeout for a single parse, or for every parse, in seconds. Passing 0 or null disables it:

Parsel::file('report.pdf')->withTimeout(120)->markdown();

Parsel::defaultTimeout(300);

Drivers

Selecting a Driver

LiteParse is the default driver. You may select another driver per parse, or change the default:

$markdown = Parsel::driver('anydoc')->file('report.docx')->markdown();

Parsel::defaultDriver('anydoc');

If you prefer dependency injection over the static facade, you may use a ParselManager instance:

use Shipfastlabs\Parsel\ParselManager;

$parsel = new ParselManager;

$markdown = $parsel->driver('anydoc')->file('report.docx')->markdown();

Capabilities

Capability LiteParse AnyDoc
Markdown Yes Yes
Plain text Yes No
Structured documents and JSON Yes No
Lazy pages Yes No
Screenshots Yes No
Local OCR Yes No

Calling an unsupported method throws an UnsupportedCapabilityException before the parser runs.

Supported Formats

Input LiteParse AnyDoc
PDF Yes Yes
Word, Excel, PowerPoint Via LibreOffice Yes
OpenDocument, RTF, CSV Via LibreOffice Yes
EPUB No Yes
Images Via OCR No

Binary Resolution

Each driver locates its executable in the following order:

  1. The binary provider option, set via withBinary().
  2. The PARSEL_LITEPARSE_BINARY or PARSEL_ANYDOC_BINARY environment variable.
  3. lit or anydoc on your PATH.

Provider Options

Driver-specific settings are passed to withProviderOptions using a typed options object or an array.

LiteParse Options

use Shipfastlabs\Parsel\Enums\ImageMode;
use Shipfastlabs\Parsel\Options\LiteParseOptions;

$markdown = Parsel::file('invoice.pdf')
    ->withProviderOptions(
        LiteParseOptions::make()
            ->pageRange(1, 5)
            ->page(10)
            ->maxPages(20)
            ->withDpi(300)
            ->withPassword($password)
            ->withImages(ImageMode::Embed, '/path/to/images')
            ->withoutLinks()
            ->keepHeadersAndFooters()
    )
    ->markdown();

OCR

You may enable OCR with withOcr(), optionally choosing a language, worker count, local Tesseract data directory, or an HTTP OCR server:

LiteParseOptions::make()->withOcr(language: 'eng', workers: 4);

LiteParseOptions::make()->withOcr(tessdataPath: '/usr/share/tessdata');

LiteParseOptions::make()
    ->withOcr(serverUrl: 'https://ocr.example.com', headers: ['Authorization' => "Bearer {$token}"])
    ->withOcrServerHeader('X-Tenant', 'acme');

[!NOTE] OCR is disabled unless you enable it, so scanned PDFs and images return empty text by default.

Structured Output

You may ask LiteParse to enrich structured output with extractBlocks, extractAnnotations, extractFormFields, extractStructureTree, extractContentBounds, extractVectorGraphics, extractTextMetadata, extractImages, extractXfaPackets, and withComplexity, or enable all of them with extractAll. These options apply to parse, toArray, lazyPages, and .json saves:

LiteParseOptions::make()->extractBlocks()->extractFormFields();

Other Options

You may skip damaged pages, load a LiteParse config file, or pass any CLI flag that Parsel does not cover yet. The option method applies to parsing, while screenshotOption applies to screenshots:

LiteParseOptions::make()
    ->continueOnPageError()
    ->withConfig('/path/to/liteparse.json')
    ->option('new-flag', 42)
    ->screenshotOption('new-screenshot-flag');

AnyDoc Options

You may set the input format explicitly or pass any CLI flag via option:

use Shipfastlabs\Parsel\Options\AnyDocOptions;

$markdown = Parsel::driver('anydoc')
    ->file('data.csv')
    ->withProviderOptions(AnyDocOptions::make()->format('csv'))
    ->markdown();

AnyDoc does not run OCR locally. You may send scanned PDFs to Firecrawl for hosted OCR instead. When no API key or URL is given, AnyDoc reads FIRECRAWL_API_KEY and FIRECRAWL_API_URL:

$markdown = Parsel::driver('anydoc')
    ->file('scan.pdf')
    ->withProviderOptions(AnyDocOptions::make()->withHostedOcr())
    ->markdown();

[!WARNING] Hosted OCR uploads the entire document to Firecrawl or the server at api_url. Only enable it for documents you are allowed to share with that service.

Array Options

Every option has a snake_case array key. Keys are validated, so typos throw an InvalidProviderOptionsException:

Parsel::file('receipt.png')
    ->withProviderOptions(['ocr' => true, 'ocr_language' => 'eng'])
    ->text();

Handling Errors

Parser, filesystem, and option failures extend Shipfastlabs\Parsel\Exceptions\ParselException. Invalid arguments, such as an empty path, an unsafe byte extension, or an unknown image mode, throw a standard InvalidArgumentException or ValueError. When a parser exits with an error, Parsel throws a ParseFailedException exposing exitCode, stderr, and command. When it exceeds the timeout, Parsel throws a ParseTimedOutException exposing timeout, driver, and command:

use Shipfastlabs\Parsel\Exceptions\ParseFailedException;
use Shipfastlabs\Parsel\Exceptions\ParseTimedOutException;

try {
    $markdown = Parsel::file('report.pdf')->markdown();
} catch (ParseTimedOutException $e) {
    report("{$e->driver} timed out after {$e->timeout}s");
} catch (ParseFailedException $e) {
    report($e->stderr);
}

The AnyDoc driver throws more specific subclasses of ParseFailedException: a ParserUsageException for invalid arguments, and an OcrRequiredException when a PDF needs OCR:

use Shipfastlabs\Parsel\Exceptions\OcrRequiredException;
use Shipfastlabs\Parsel\Options\LiteParseOptions;

try {
    $markdown = Parsel::driver('anydoc')->file('scan.pdf')->markdown();
} catch (OcrRequiredException) {
    $markdown = Parsel::file('scan.pdf')
        ->withProviderOptions(LiteParseOptions::make()->withOcr())
        ->markdown();
}

Secrets such as passwords, API keys, and OCR server headers are redacted from the exception's command, so it is safe to log.

Custom Drivers

You may register your own driver with the extend method. A driver implements the Driver contract, which provides Markdown, and may implement TextDriver, StructuredDocumentDriver, LazyPageDriver, or ScreenshotDriver for additional capabilities:

use Shipfastlabs\Parsel\Contracts\Driver;
use Shipfastlabs\Parsel\ParselManager;

Parsel::extend('company-api', fn (ParselManager $manager): Driver => new CompanyApiDriver);

$markdown = Parsel::driver('company-api')->file('report.pdf')->markdown();

The factory receives the ParselManager. Its processRunner and filesystem methods return the process runner and filesystem the bundled drivers use, so a driver built on them respects Parsel::fake().

When withProviderOptions is called more than once, Parsel merges string pages selections and the extra and screenshot_extra arrays. Every other key is replaced by the latest value.

Testing

The fake method replaces the parser processes with canned responses, matched against a substring of the command:

$fake = Parsel::fake([
    '--format json' => file_get_contents(__DIR__.'/fixtures/lit-output.json'),
    'anydoc' => '# Converted document',
]);

$document = Parsel::bytes($pdf, 'pdf')->parse();
$markdown = Parsel::driver('anydoc')->bytes($docx, 'docx')->markdown();

expect($fake->ranCount())->toBe(2);

Only the parser processes are faked, so a path passed to file must still exist. Use bytes when the document isn't on disk.

The fake, default driver, default timeout, and registered drivers are kept for the whole process. Call Parsel::flush() in your test teardown to reset them:

afterEach(fn () => Parsel::flush());

You may return a ProcessResult to simulate a failure or timeout:

use Shipfastlabs\Parsel\Support\ProcessResult;

Parsel::fake(['anydoc' => new ProcessResult(143, '', '', ['anydoc'], timedOutAfter: 30.0)]);

Upgrading

Please see UPGRADE.md for upgrade instructions.

Contributing

You may run the test suite with Composer:

composer test

The integration tests run against the real lit and anydoc binaries. Set PARSEL_REQUIRE_BINARIES=1 to fail instead of skip when a binary is missing, and PARSEL_INTEGRATION_EXTENDED=1 to include tests that need network access or LibreOffice:

vendor/bin/pest --group=integration

Credits and License

Parsel is maintained by Shipfastlabs and released under the MIT license.

More from the ecosystem

marcreichel/
laya-php
g4t.io
3 53

LayaPHP: Self-Hosted Text Classification for PHP and Laravel

Classify text in PHP without an LLM bill: typed decisions in 100+ languages, self-hosted. Laravel-ready SDK for Laya, a Jev AI alternative.

ai classification decision-engine
marcreichel/laya-php via Laravel News
Josh-Dovey/
postcodes-laravel
g4t.io
5 9

Postcodes for Laravel: GB Postcode Lookup and Geography Data

Postcodes for Laravel adds typed GB postcode lookups, validation, geography data, distance searches, and test fakes through the GB Postcodes API.

Josh-Dovey/postcodes-laravel via Laravel News
RedberryProducts/
mailbox-for-laravel
g4t.io
4 115

Mailbox for Laravel: Preview and Test Rendered Email

Mailbox for Laravel captures outgoing mail in a local dashboard and lets you test rendered HTML, recipients, and attachments with fluent assertions.

RedberryProducts/mailbox-for-laravel via Laravel News