Documents and versions
A document in a house is not a file. It is an identity whose content is a structure of blocks: sections, paragraphs, lists, quotes, code. The calls below were run against a fresh house.
Writing
create_documentwith a name creates the identity. Every call creates a new one; names are labels, and creating never searches. Checklist_documentsfirst if reuse might be intended.import_documentwith the document id and Markdown asserts the whole state. The answer sayschanged: true, gives the id of the new version root and the version count.
Importing the identical Markdown again is a recognised no-op (changed: false). Importing
a changed text makes a new version: in the test a second import with one changed paragraph
answered version_count: 2. Blocks that did not change are literally the same nodes in
both versions.
write_document is the structured twin of import_document for editors that work on the
block tree instead of Markdown.
Reading
export_documentreturns the canonical Markdown of the current version, or of any earlier one by its version root id. It returns a character window (4000 by default, up to 30000 per call) with the total length and the offset to continue at.read_documentreturns the block tree: every block with its id and its form.list_documentslists all documents with name, version count and current root.- The versions of a document hang in its
Structureslot.get_historyreads them, and needs the Document aspect as context. If you forget it, the house tells you exactly what to pass: the error names the slot and the context id.
Searching
The text of a document is indexed as it arrives. search_text for "stone bridge" returned
the paragraph, and with it the document it stands in and whether that version is the
current one, so no follow-up call was needed. find_occurrences answers for any text in
which documents and versions it appears.
Taking material in from outside
import_binarytakes bytes into the house. The bytes travel on a separate channel:POST /binaryon the house's address returns their content hash, and the skill takes the hash.GET /binary/{hash}returns them.import_capturetakes a page captured in a browser (an MHTML archive or an HTML fragment uploaded the same way) and reads a cleaned document out of it.import_urllets the house fetch a web page itself and keep both the original bytes and a cleaned document.
What is read out of the bytes depends on the way in. A capture and a fetched page become a
document at once: in the test an HTML fragment came back as a document with two blocks. A
file taken in with import_binary is read into a document when you pass its mime and the
house has a reader for it: PDF with a text layer, DOCX, PPTX, DOC, Markdown, plain text and
HTML. The answer names the document, or says why nothing was read; without mime the bytes
are only stored. In the sandbox, on 2026-10-07, a DOCX and a PDF each came back as a document,
and search_text found their words a few seconds later. Text from a PDF or a plain text file
has lost its layout; putting headings and paragraphs back is a judgment that waits on the
list of open work (the-list-of-open-work). OCR for scans does not exist yet: a scan is
stored, not read.
What a house does not do by itself
A house does not classify, summarise or outline a document on its own. Those are
judgements; see who-does-the-thinking.