Cutting a skill in half without making it worse
Shortening two caliper skills, grill-skill and evaluate-skill, by ~45% using a non-inferiority eval bar: prove the shorter version is no worse before shipping the cut.
AI safety notes
Essays, notes, and working models for AI safety, agentic systems, and trustworthy infrastructure.
Shortening two caliper skills, grill-skill and evaluate-skill, by ~45% using a non-inferiority eval bar: prove the shorter version is no worse before shipping the cut.
A case study in using Caliper to evaluate blader/humanizer, tighten voice calibration, and turn the improvement into an upstream contribution with regression coverage.
Tribal knowledge encoded as an AI skill is still just text until you evaluate it. Ablation baselines, routing regression tests, trajectory autoraters, and the gotchas flywheel keep encoded knowledge from rotting.
Everyone is worried about AI reading things it shouldn't. That's the wrong threat model. The problem starts after the agent reads.
Skills bundle instructions, scripts, and MCP servers into a single installable package. That convenience is also the attack surface.
A few words on what this space is about.