Our approach to safety and safeguards
Private AI still needs guardrails, but we should not have to read your conversations to enforce them. Our policy and code are public, and remote attestation lets you verify that the code running inside each enclave is exactly what we publish.
Evaluate before hosting
Every chat model is tested against our hard-no policy before we host it. The prompts, grading criteria, and full model answers are open source, so you can check our results rather than take our word for them.
Review inside our enclaves
Safeguards run inside our secure enclaves. If a conversation violates the policy, only a flag and the conversation ID leave the enclaves. The conversation itself, and which rule it broke, never reach us.
Transparent policy
The hard-no policy covers three things: child endangerment, mass violence and terrorism, and self-harm. Nothing else is monitored. Our terms of service define broader acceptable use.
Safety evaluation methodology
What we test
Before hosting a model, we evaluate its responses to handpicked prompts that match our hard-no policy.
How to read the numbers
Pass rate is the share of prompts handled correctly. Category percentages show how failed responses are distributed, not failure rates within each category.
What it does not cover
These single-turn benchmarks cannot cover every conversation or jailbreak. Our active safeguards provide an additional check during chat.
Open source
Our methodology, prompts, and full model answers are public, so you can rerun any evaluation yourself.
View benchmarksTinfoil provides a set of third-party open-source models and gives you the freedom to choose which model to use based on your own evaluation. To help you choose the right model, we provide information about each model and how it met our listing standards, but the decision and responsibility remain yours.