Recipe for assigning documents to topics

When to use it

Use it when a house has topics and documents that are not yet sorted under them. It is the first recipe in the order of working-by-recipe.

Use this recipe on documents that arrive with structure (Markdown, or HTML from a web capture). Documents that arrive as raw text or from PDF need their outline first, from the list of open work (list_work): an outline creates a new version of the document, and work done on the old version will be offered again.

What you need first

Topics with descriptions. overview must list an aspect named Thema, and list_forms a form titled About. If either is missing, or the owner has no topics yet, follow topics-in-a-house first. Topics are the owner's: never coin one because a document fits nowhere.

Fetch the context

Once per run:

overview {}
list_collection {"name": "<instance_collection of Thema>", "limit": 200}
get_property {"thing_id": "<topic ID>", "path": ["Umschreibung"], "context_id": "<Thema concept ID>"}
list_forms {}
list_documents {"limit": 200}

The answer of overview is large. Take the entry of aspects whose name is Thema: its concept.id is the concept ID, and its instance_collection is the name to pass to list_collection, exactly as given (it is not the concept ID). The topics are the items. get_property is called once per topic; slots[0].text is the description. Sort the topics by name in plain code-point order and number them from 1. From list_forms take form.id of the entry whose title is About. list_documents pages with offset; skip digests, as working-by-recipe says.

Per document:

get_links {"id": "<document ID>", "direction": "out", "form_id": "<About ID>"}
get_property {"thing_id": "<document ID>", "path": ["TopicsClaimed"]}
export_document {"doc_id": "<document ID>", "max_chars": 30000}

The first shows the assignments that stand. The second shows the topics ever asserted at this document; where none were, slots is empty and a note speaks of a missing context, which is no error. One judgment reads at most 40,000 characters: if total_length is larger, take the first 30,000 and the last 10,000 (a second export_document with offset) and put the line [… N characters left out …] between them. Naming the gap keeps a model from inventing into it.

Build the input exactly so. The title is the name from list_documents. The export goes in as it is, even where it begins with the title again:

# Topics

1. **Bridge renovation**
   Everything about inspecting and repairing the harbour bridge. Not the trust's finances.

2. **Money**
   (no description — judge by the name alone)

# Document

Title: Budget note, second quarter

<the Markdown of the document>

Judge

One judgment per document, with this prompt:

You assign a find to the standing topics of the owner of this memory — pages that were
saved, documents that were imported, notes that were written.

You get the list of topics (each with a description the owner wrote) and one document.
Say which topics the document belongs to.

How to judge:
- The description of a topic decides, nothing else. It says what belongs there.
- Be generous: a document may belong to several topics, and an assignment that only
  partly holds is better than none. Too much is easy to correct, too little stays
  invisible.
- Still force nothing: if the document fits no topic, return an empty list. That is a
  valid result, not a failure.
- Give each assignment ONE concrete sentence that names the content — "Interview about
  how freelancers set their prices", not "fits the topic". Write it in the language of
  the topic descriptions. The sentence is later shown to the owner as the explanation.

Answer only with JSON of the agreed form: {"assignments": [{"topic": <number from the
list>, "reason": "<one sentence>"}]}.

The answer form

{"assignments": [{"topic": 1, "reason": "Report on corroded bearings and their replacement."}]}

topic is the number in your list, an integer. An empty list means "no topic".

Check your own verdict

Apply

Assert first, then record what you asserted:

assert_formed_link {"form_id": "<About ID>", "from_id": "<document ID>", "to_id": "<topic ID>",
                    "values": {"Reason": "<the sentence>"}}
set_property {"thing_id": "<document ID>", "path": ["TopicsClaimed"],
              "value": "claimed-v1 <topic ID>,<topic ID>", "versioned": true}

One assert_formed_link per assignment; is_new: false means it stood already. The value of TopicsClaimed is the prefix claimed-v1, a space, and all topic IDs ever asserted at this document, sorted in plain code-point order and joined by commas: those read before plus those of this verdict, new or not. The prefix matters: a value that is nothing but an ID would be stored as a link to that topic. Skip the call when the list is empty. The owner corrects an assignment with retract; you never retract one.

Say who judged

Leave one note per document, as working-by-recipe shows, naming the recipe and the topics you assigned, or that none fitted.

What this recipe does not do

Origin

Atlantis, Intelligence/World/Topics.Assigning/TopicJudge.cs (prompt, answer form, input) and Assigner.cs (apply), commit 5135ad2, read at main 966157b. Adapted: the prompt ran in German and spoke of one person's finds; this is its English form, word for word that of assign_topics on the list of open work in Atlantis, Core/Interface/Skills/Working/TopicWork.cs, commit 99ce970, with the rules of judgment unchanged.

Go deeper