LLM reasoning / chain-of-thought

H-CoT

H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking

Superseded baseline#11 of 772 most-superseded · first seen Feb 18, 2025

Superseded — cited as a baseline and beaten by newer methods

2 papers critique it · 2 beat it on benchmarks

What papers say

Verbatim critique sentences, each from a paper that cites H-CoT as a baseline.

it remains ineffective on the latest o3 and o4-Mini
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
While effective in some cases, such approaches suffer from three major limitations. First, their reliance on fixed templates restricts diversity, making attacks easier to detect or defend against. Second, they lack adaptability to different models and contexts, limiting their robustness. Third, their overall effectiveness is constrained, as static designs fail to fully exploit the dynamic nature of CoT reasoning.
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs

Beaten on benchmarks

Head-to-head results where a newer method reports beating H-CoT. Values are copied from the source paper's tables — verify against the cited paper.

What to use instead

Recent methods in the same sub-problem, not yet superseded in the knowledge base — arXiv benchmark leaders, not vetted production recommendations.