SHARE
Empirical Inference Conference Paper 2026

Training with Honeypots: Reshaping How LLMs Fail

Empirical Inference
Research Scientist
Empirical Inference
Author(s): Simko, S. and Pandey, P. S. and Jin, Z. and Schölkopf, B.
Book Title: ICLR 2026 Workshop on Principled Design for Trustworthy AI - Interpretability, Robustness, and Safety across Modalities
Year: 2026
Month: April
BibTeX Type: Conference Paper (conference)
Event Name: ICLR 2026 Trustworthy AI
Event Place: Rio de Janeiro, Brazil
State: Published
URL: https://openreview.net/forum?id=yP24gVeeFo

BibTeX

@conference{SimPanJinSch26,
  title = {Training with Honeypots: Reshaping How LLMs Fail},
  booktitle = {ICLR 2026 Workshop on Principled Design for Trustworthy AI - Interpretability, Robustness, and Safety across Modalities},
  month = apr,
  year = {2026},
  author = {Simko, S. and Pandey, P. S. and Jin, Z. and Sch{\"o}lkopf, B.},
  url = {https://openreview.net/forum?id=yP24gVeeFo},
  month_numeric = {4}
}