Tagged: evals
11 articles tagged evals
- How many AI agents should review your code? I tried 99.
- Do AI engines cite my site? Measuring it without inventing an AI rank
- The local model wasn't worse. My context pipeline was.
- Is a free local AI model actually cheaper? Mine wasn't
- Why was the check green while ten records were missing?
- A URL in the fetch log is not a citation
- The AI said the workbook was done. Ten of eleven tabs were empty.
- Before an AI agent delegates work, can it tell which tasks are hard?
- I gave my AI tests it didn't know were tests. Here is its real score.
- My AI designed, ran, graded — and won — its own benchmark. I threw the run out.
- Before locking the cuts, I ordered the AI to argue against itself