Bias and Edge-Case Evaluation Set
A curated evaluation set for probing AI consistency, fairness behaviour and edge cases. Designed for testing, not training.
Intended use
- Validation
- Red-team
- Demonstration
Not intended for
- Training
- Claims of full population fairness
Documentation level
High
Overview
Carefully constructed evaluation scenarios for probing model behaviour on sensitive prompts, ambiguous instructions, contradictory contexts and known failure modes.
What it contains
- Dataset file (CSV)
- Sample preview (100 records)
- Data dictionary
- waiser data Dossier
- License terms
- Version log
Generation method
Curated by hand from published red-team literature and adapted scenario templates. Each item is reviewed for clarity and scope.
Permitted and prohibited uses
Permitted
- Model evaluation
- Internal red-team exercises
- Demonstration of evaluation pipelines
Prohibited
- Training
- Public claims of comprehensive fairness assessment
- Use outside the permitted license scope
Dataset structure
| Column | Description |
|---|---|
| item_id | Unique item identifier |
| category | Bias, ambiguity, contradiction, edge_case |
| prompt | Evaluation prompt |
| expected_behaviour | Description of expected model behaviour |
| severity | Low / medium / high |
Quality notes
Coverage is limited to documented categories. Not a substitute for full fairness audit. Marked v0.9 — community feedback welcome.
waiser data Readiness Profile
Ethical notes
Some items reference sensitive topics by design. Reviewers should be briefed in advance and provided context for interpretation.
License
waiser data Standard License — Defined Use
Full license terms are provided with the dataset delivery. Access is granted after review of intended use.
Version log
ready to request access?
Tell us your intended use and we'll respond with the executed licence, the full dossier and a secure delivery link bound to your request.