Practice the distinctions.
Choose an answer to see the reason behind it. Reveal a card after trying to explain it yourself. Attempts are saved locally, without points or rankings.
The AI claim detective
What is the strongest test of your prompt?
Fluency is not evidence
A model predicts plausible text. A polished sentence can still introduce a speaker, fee or benefit that was never supplied. Give every factual sentence an evidence line or mark it unknown.
A prompt is a contract
Specify the job, approved evidence, constraints and output shape. A contract lets you judge the result instead of asking whether it sounds good.
Evaluation starts small
Use a normal request, a missing fact, conflicting facts and an instruction that asks the model to invent details. Track which rule each output passes or fails.
Weekend Quest: an assistant with boundaries
Where should current availability go?
The assistant says it booked a room, but has no booking tool. What do you do?
Role is different from memory
A role card describes the job and its limits. Current facts belong in the request. Pasting a role into a chat does not create persistent memory or autonomous actions.
Constraints need priority
A required indoor location overrides a preference for gardening. Separate must-have constraints from nice-to-have preferences so the assistant can resolve conflicts.
A useful plan is inspectable
A plan should show its evidence, duration and unanswered questions. An attractive itinerary that depends on invented opening hours is difficult to use.
Neighbourhood Swap: your first useful website
What is the real deliverable?
Which change belongs mainly to CSS?
HTML carries meaning
Headings, lists, navigation and buttons tell browsers and assistive tools what content does. A visual div is not automatically an accessible control.
CSS controls presentation
Spacing and responsive rules should adapt the same content to different screens. Content should remain readable when the viewport changes or text is enlarged.
A website is an artifact
A prompt is a request for a file. The deliverable is the saved code that opens in a browser and survives your tests, not a screenshot of a chat answer.
BorrowBox: turn a page into a working app
Why can a second phone see an empty list?
How should item names be rendered?
State is the source of truth
Keep records in an array with stable IDs. Render the screen from those records. Reading the current DOM as your only database makes changes difficult to reason about.
Persistence is separate from display
localStorage can save a JSON string for this browser origin. It does not create a shared community database or a cross-device backup.
Inputs are untrusted text
Trim empty names, prevent duplicates according to your rule and render names using textContent. Treat an entered HTML tag as text, not executable markup.
Pocket BorrowBox: put your app on a phone
What does a service worker cache guarantee?
Why test a cache update?
A PWA is a web app
A manifest supplies app identity and install information. Browser support and installation UX vary, so test the device rather than assuming every phone shows the same prompt.
Cache is not data sync
A service worker can cache static app files. It does not automatically synchronize records, run a remote AI model offline or turn localStorage into cloud storage.
Versions matter
An old cached file can keep showing an outdated app. Name cache versions, remove obsolete caches in activation and test an update before sharing.
Learning Radar: a second brain that explains what to learn next
What does Obsidian Graph View do?
An archived note says a standard bot can initiate a private chat. Current documentation contradicts it. What should you do?
Links express relationships
In Obsidian, notes and internal links form a graph. The graph makes relationships visible; it does not itself reason or call an AI. Our animated crawler is a teaching simulation of evidence selection.
Retrieval comes before an answer
First find the notes relevant to a question, then use those notes as the answer context. Providing every note can bury useful evidence or include stale facts.
A source is more than a filename
Record what the note says, when it was updated and whether it conflicts with another note. Prefer verified current evidence and explain the conflict instead of blending incompatible facts.
BorrowBox dispatcher: a reliable automation
Why use fixed conditions here?
Which branch runs first?
Rules can be enough
A known routing policy is easier to inspect as explicit conditions. An LLM would introduce variability without adding a needed capability to this task.
Validate before routing
Missing item IDs and strings pretending to be booleans can produce incorrect branches. Check shape and type before making the routing decision.
Retries need identities
If a trigger repeats a request, a stable request ID lets you recognize duplication. Otherwise a reliable retry can accidentally create two jobs.
Meetup inbox: AI that produces checkable data
Valid JSON always means a reliable answer.
What should happen to conflicting day preferences?
A chain is a fixed model step
The workflow chooses when the model runs. The model drafts or transforms text; it does not choose and invoke external tools in this basic chain.
Structured output is a contract
A JSON-shaped answer is not automatically valid JSON or valid data. Parse it and validate required fields, types and evidence references before downstream actions.
Untrusted input stays data
An incoming message saying ignore your rules is part of the message content. It cannot authorize sending, changing the workflow or inventing new facts.
A neighbourhood concierge bot
What identifies the reply destination?
Why use a dedicated test bot?
A bot has an event boundary
A Telegram message creates an update. A trigger receives it, a workflow chooses a reply and the Telegram API sends the response. A text prompt alone is not a deployed bot.
One webhook needs a clear owner
Test and production configurations can compete for a bot webhook. Use a dedicated test bot and understand which workflow currently receives its updates.
Unknown commands need a route
A resident should get useful help when a command is absent or unsupported. Silence and invented answers both make a bot hard to trust.
The setup planner: an agent you can inspect
What proves calculator use?
A weather question arrives with only Calculator connected.
Tool choice is the distinguishing step
The model can choose a connected tool and use its result. A fixed chain has a predetermined model call instead. Merely writing the word agent in a prompt does not connect capabilities.
Tool output is evidence
The arithmetic result should come from the calculator call and be visible in the trace. A plausible narrated tool call is not proof that an execution happened.
A stopping rule makes failures inspectable
Set a maximum iteration count and an error path. If the task needs an unavailable tool, the agent should explain the missing capability rather than inventing a result.
A writing twin with an honest identity
Does matching someone’s style make facts true?
Which style instruction is easier to evaluate?
Style and facts are separate inputs
A writing example teaches tone, sentence length and structure. It does not establish that every fact in a new answer is true. Supply approved facts independently.
A twin is a draft assistant
This exercise creates text that follows a style profile. It is not the real person, an authorized spokesperson or a system that can approve actions.
Good comparisons use a rubric
Compare the same request with and without the style profile. Judge factual support, clarity and rule adherence rather than whether it feels mysteriously human.
Neighbourhood HQ: your personal assistant cockpit
An approval field is missing. What should happen?
When should you add another agent?
Composition needs contracts
Each component should have explicit input and output. A notes step returns evidence; a model step returns a draft; an approval step returns a decision. Ambiguous handoffs make failures hard to locate.
Approval is a state transition
A request to approve is not approval. Keep pending, approved and rejected states separate; only the approved branch can save or send the selected plan.
Observability enables independence
Record a run ID, evidence IDs, validation result, approval outcome and final artifact. When something fails, diagnose the component instead of blindly changing the whole prompt.