Flowers
An open-source personal-agent harness, built to try out reliability ideas. It plans before it acts and asks before anything irreversible.
Testbeds, side systems, and papers outside our main line of work. Two of the projects run in your browser.
An open-source personal-agent harness, built to try out reliability ideas. It plans before it acts and asks before anything irreversible.
Verified symbolic memory for LLMs. It checks the claims in a draft answer against Wikidata and abstains where it cannot find grounding.
An autonomous multi-agent loop that probes the structure of the Voynich Manuscript. It ran over 600 corpus-statistical experiments and over 5400 cipher-mechanism simulations.
Papers by lab members, outside our main line of work.
A companion to probe-and-refine: can comparable guidance be assembled without ever tuning on the target repo? Mining transferable edits from 50 external repositories yields guidance that resolves 45.4% of SWE-bench Verified vs. 39.6% baseline.
Read →