Agents need an operating model, not another framework
One agent asked to plan, build, test, and validate still fails on large work the way large human projects fail. Skailr is the org structure we install on top of Claude Code and Cursor.
Read →INSIGHTS
Short pieces on agents, applied ML, evaluation, and the production constraints that shape our research.
LATEST
One agent asked to plan, build, test, and validate still fails on large work the way large human projects fail. Skailr is the org structure we install on top of Claude Code and Cursor.
Read →Generic fake data looks fine in demos and breaks when you test record matching systems. Here is what entity resolution evaluation harnesses actually need from synthetic test data.
Read →Entity resolution systems are scored on the duplicates they catch and the false merges they avoid. A test set of unique, well-formed rows cannot exercise either.
Read →Test data is only useful where your pipeline can read it. Export formats and a programmatic API matter as much as the data itself.
Read →Routing, planning, and adapting across agents is only trustworthy if you can see it happen. Autonomy without observability is a black box that occasionally works.
Read →Ad-hoc scraping gives you text about the market. An ontology gives you a structure you can query, reason over, and eventually act on.
Read →Matching pipelines, scoring models, and agentic systems all fail on the same class of case: the messy one a tidy demo never shows you.
Read →