Documents and versions

A document in a house is not a file. It is an identity whose content is a structure of blocks: sections, paragraphs, lists, quotes, code. The calls below were run against a fresh house.

Writing

  1. create_document with a name creates the identity. Every call creates a new one; names are labels, and creating never searches. Check list_documents first if reuse might be intended.
  2. import_document with the document id and Markdown asserts the whole state. The answer says changed: true, gives the id of the new version root and the version count.

Importing the identical Markdown again is a recognised no-op (changed: false). Importing a changed text makes a new version: in the test a second import with one changed paragraph answered version_count: 2. Blocks that did not change are literally the same nodes in both versions.

write_document is the structured twin of import_document for editors that work on the block tree instead of Markdown.

Reading

Searching

The text of a document is indexed as it arrives. search_text for "stone bridge" returned the paragraph, and with it the document it stands in and whether that version is the current one, so no follow-up call was needed. find_occurrences answers for any text in which documents and versions it appears.

Taking material in from outside

What is read out of the bytes depends on the way in. A capture and a fetched page become a document at once: in the test an HTML fragment came back as a document with two blocks. A file taken in with import_binary is read into a document when you pass its mime and the house has a reader for it: PDF with a text layer, DOCX, PPTX, DOC, Markdown, plain text and HTML. The answer names the document, or says why nothing was read; without mime the bytes are only stored. In the sandbox, on 2026-10-07, a DOCX and a PDF each came back as a document, and search_text found their words a few seconds later. Text from a PDF or a plain text file has lost its layout; putting headings and paragraphs back is a judgment that waits on the list of open work (the-list-of-open-work). OCR for scans does not exist yet: a scan is stored, not read.

What a house does not do by itself

A house does not classify, summarise or outline a document on its own. Those are judgements; see who-does-the-thinking.

Go deeper