# Documents and versions A document in a house is not a file. It is an identity whose content is a structure of blocks: sections, paragraphs, lists, quotes, code. The calls below were run against a fresh house. ## Writing 1. `create_document` with a name creates the identity. Every call creates a new one; names are labels, and creating never searches. Check `list_documents` first if reuse might be intended. 2. `import_document` with the document id and Markdown asserts the whole state. The answer says `changed: true`, gives the id of the new version root and the version count. Importing the identical Markdown again is a recognised no-op (`changed: false`). Importing a changed text makes a new version: in the test a second import with one changed paragraph answered `version_count: 2`. Blocks that did not change are literally the same nodes in both versions. `write_document` is the structured twin of `import_document` for editors that work on the block tree instead of Markdown. ## Reading - `export_document` returns the canonical Markdown of the current version, or of any earlier one by its version root id. It returns a character window (4000 by default, up to 30000 per call) with the total length and the offset to continue at. - `read_document` returns the block tree: every block with its id and its form. - `list_documents` lists all documents with name, version count and current root. - The versions of a document hang in its `Structure` slot. `get_history` reads them, and needs the Document aspect as context. If you forget it, the house tells you exactly what to pass: the error names the slot and the context id. ## Searching The text of a document is indexed as it arrives. `search_text` for "stone bridge" returned the paragraph, and with it the document it stands in and whether that version is the current one, so no follow-up call was needed. `find_occurrences` answers for any text in which documents and versions it appears. ## Taking material in from outside - `import_binary` takes bytes into the house. The bytes travel on a separate channel: `POST /binary` on the house's address returns their content hash, and the skill takes the hash. `GET /binary/{hash}` returns them. - `import_capture` takes a page captured in a browser (an MHTML archive or an HTML fragment uploaded the same way) and reads a cleaned document out of it. - `import_url` lets the house fetch a web page itself and keep both the original bytes and a cleaned document. What is read out of the bytes depends on the way in. A capture and a fetched page become a document at once: in the test an HTML fragment came back as a document with two blocks. A file taken in with `import_binary` is stored and named, nothing more. The house has extractors for PDF and office formats, and `coverage` lists them per format, but through the skills they are not yet run: a PDF taken in this way showed as `untouched`. Until that changes, give the house the text as Markdown with `import_document` and keep the original as a binary beside it. OCR for scans does not exist yet. ## What a house does not do by itself A house does not classify, summarise or outline a document on its own. Those are judgements; see `who-does-the-thinking`. ## Go deeper - [finding-things](finding-things.md) - [who-does-the-thinking](who-does-the-thinking.md) - [evidence-and-provenance](evidence-and-provenance.md) - [the-building-blocks](the-building-blocks.md)