Skip to main content
Before you can search, your documents need to be in LightOn. Uploading a file triggers an ingestion pipeline that parses the content, splits it into chunks, generates embeddings, and indexes everything. The whole process typically takes a few seconds for a standard PDF. Ingestion is asynchronous: the upload returns immediately with a pending status, and you poll GET /api/v3/files/{id} for completion.
This tutorial covers POST /api/v3/files and GET /api/v3/files. The full schema for every endpoint and parameter lives in the API reference.

Upload a file

Send the file as multipart/form-data with the destination workspace_id:
The response is a 201 with the new file record, including an upload_session_uuid you can use later to find every file uploaded in the same batch.

Wait for indexing to complete

Poll GET /api/v3/files/{id} until status reaches embedded:
The status field moves through these stages: status_vision tracks the same lifecycle for vision/image embeddings: pending, processing, embedded, fail, or - (not available for this file). One caveat: status describes the run that has already happened. If reprocessing is queued but hasn’t started, pending_reprocess is non-null and status still reports the previous run. That only comes up after a file replacement, where trusting status alone would report success against the old file.

Organising documents with tags and titles

Add a human-readable title and assign tag IDs at upload time. Tags can be sent as a JSON-encoded array string or as repeated form fields with the same name.
If a tag ID is invalid, the file is still created but the response is a 207 (multi-status) with a message explaining which tags were rejected. To replace tags after upload, PATCH /api/v3/files/{id} with a new tags array. It replaces all existing tags, manual and auto-assigned. Send [0] (sentinel) to remove every tag when using multipart format. To add tags without touching existing ones, POST /api/v3/files/{id}/tags.

Tracking documents from external systems

If you’re ingesting documents from a third-party system (ServiceNow, Confluence, SharePoint, etc.), store the source identifier in external_metadata. This lets you find the LightOn file later given only the external ID, and surface the original URL in your UI.
external_id is required when creating; doc_type and additional_metadata are optional. When sent via multipart/form-data, the whole external_metadata value must be a JSON string. Updates merge rather than replace, including inside additional_metadata, so you can patch one key and leave the rest alone. external_id cannot be changed on a document imported from a datasource: the sync matches its documents on it, so a rename would detach the document from its source. Retrieve it later by external ID:

Listing and filtering your documents

GET /api/v3/files supports rich filtering. A few common patterns:
Set include_details=true to receive the signature (TLSH hash for duplicate detection) and parser fields on each result. For advanced metadata filtering, see the Facets tutorials: classify files by type, set custom attributes, then filter with operators. See the filter reference for the full DSL.

Downloading a document

Once a file is ingested, GET /api/v3/files/{id}/download gives you the stored bytes back, so your app can offer the original document next to the answer it grounded.
The purpose parameter picks which stored version you get: If the requested purpose has no associated file, the API falls back to original rather than returning a 404.

Thumbnails

Every document also gets a 256x256 WebP thumbnail, handy for a document picker or a result list. Generation runs asynchronously and separately from ingestion, so an embedded file may still have none. Check the thumbnail.status field on the file before fetching, or GET /api/v3/files/{id}/thumbnail returns a 404:
status is MISSING, PROCESSING, or READY. url is populated only once it reads READY.

Reading a document’s parsed text

The platform already parsed your document at ingestion, so you don’t need to re-upload it to /parse to read the text. Pass include_content=true to the file detail endpoint and you get the same {index, markdown} page shape /parse returns:
The text can be large, so include_content is off by default. Ask for it only when you need the content, not on every list call.

Replacing a document’s file

When the source document changes, you don’t have to delete and re-upload: send the new file as the file part of PATCH /api/v3/files/{id} and the document is re-ingested in place. It keeps its id, title, manual tags, content types and attribute values, so every reference you stored still resolves.
Only the file and what is derived from it (text, chunks, embeddings, auto-assigned tags, summaries) is rebuilt. The new file may be of a different type than the one it replaces. You can send title, tags or external_metadata in the same request: the metadata changes apply immediately, the file is queued.
The response describes the document as it still is. status typically reads embedded (the previous run), and filename, size, total_pages and summaries still describe the old file. The one field that tells you a replacement is pending is pending_reprocess, which reads update until the pipeline picks the document up.So poll until pending_reprocess is null and status is terminal. Treating a terminal status on its own as completion reports success against the previous file. The SDK’s wait() already accounts for this, which is what replace(..., wait=True) uses.
The document keeps serving its previous content until processing actually starts, so search, extract and download stay consistent in the meantime. Once conversion begins, a document whose new file needs rendering has no rendered PDF yet, and download?purpose=rendered_pdf serves the original instead. Gate an in-browser viewer on status, not on the download succeeding. Replacing a file needs an owner or editor role in the document’s workspace, and is rejected in workspaces configured for synced documents. The storage quota is checked on the size difference, not the full new size.

Deleting files

Single delete:
Bulk delete:
Both return 204 No Content on success. Files in synced (datasource-managed) workspaces cannot be deleted manually. The API returns 400.

Common errors