<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>securesein — Evaluation</title><description>Benchmark design, contamination, harnesses, construct validity.</description><link>https://securesein.com/</link><item><title>How Do You Catch a Machine in a Lie?</title><link>https://securesein.com/blog/catching-ai-liars/</link><guid isPermaLink="true">https://securesein.com/blog/catching-ai-liars/</guid><description>Nineteen teams spent a month building AI lie detectors. The winning trick turned out to be less clever than it looked — and that&apos;s the interesting part.</description><pubDate>Sat, 12 Sep 2026 00:00:00 GMT</pubDate><author>Sebastiaan with AI</author><category>Research — Deep dive</category><category>ai-safety</category><category>evaluation</category><category>llms</category></item><item><title>LLMs Model a Harsher World Than the One We Live In</title><link>https://securesein.com/blog/llms-model-a-harsher-world/</link><guid isPermaLink="true">https://securesein.com/blog/llms-model-a-harsher-world/</guid><description>A new benchmark shows language models know what&apos;s wrong but not how people actually react to it — and the bias may come from alignment itself.</description><pubDate>Wed, 09 Sep 2026 00:00:00 GMT</pubDate><author>Sebastiaan with AI</author><category>Research — Paper</category><category>llms</category><category>ai-safety</category><category>evaluation</category></item></channel></rss>