Renforce
Seven ways to watch something learn.
Most explanations of reinforcement learning are diagrams. These are products — a desk, a town, a poker table, a market with two bots in it — and each one makes a claim out loud and ships a command that tries to break it. Where the mathematics has a closed-form answer, that is what the claim is checked against, rather than a baseline.
7projects on one floor
0of them running
0new dependencies to build them
The floor, which is built
Four of the seven need gradients and all seven need to say how sure they are, so those two things exist before any of the projects do. pnpm renforce:prove checks them and exits non-zero if any claim breaks.