Cross-dataset generalisation is a fundamental requirement for deploying text classifiers in real-world settings, yet systematic evaluation across corpora from different sources remains uncommon in fake news detection and virtually absent in sarcasm detection research. This paper presents a unified empirical study of zero-shot cross-dataset transfer in three domains: Urdu fake news detection (FND), English FND, and sarcasm detection. For each domain, we fine-tune xlm-roberta-base on one corpus and evaluate it on a second corpus from a different source, comparing against TF-IDF baselines with Logistic Regression (LR) and Support Vector Machines (SVM). In Urdu FND (Ax-to-Grind vs. Notri-Fact), we identify a severe length confound in the Ax-to-Grind dataset, fake articles average 3.4 times more words than real articles, causing catastrophic A to B transfer collapse (macro F1 = 0.005) while B to A achieves F1 = 0.771. Extension to English FND (WELFake vs. ISOT) and sarcasm detection (TweetEval Irony vs. Sarcasm Corpus V2) reveals that such failure modes extend beyond Urdu, confirming that shortcut learning from distributional artefacts is a cross-lingual, cross-domain challenge in binary text classification. We provide a reusable diagnostic methodology, combining class-conditional length analysis, bidirectional transfer asymmetry, and predicted label collapse inspection, applicable across any binary text classification setting.
Joe Watson, Joana Ribeiro de Faria, Marcus Tomalin +6cs.AI
Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting. This paper empirically investigates shortcut learning in this context by studying claim-level outcome prediction in UK Employment Tribunal (UKET) decisions. Using a corpus of 33,158 individual claims, we predict outcomes from claim texts and LLM-extracted case summaries, evaluating models ranging from interpretable TF-IDF-based classifiers to black-box LLMs. While headline predictive performance figures appear strong, we demonstrate that such performance in LJP systems trained on post-hoc judicial text can be driven by the retrospective nature of the source material. Stratifying the test data by human judgments of leakage reveals that performance increases where outcome-revealing cues are embedded in the narrative. Moreover, a model trained on just the 4% of features identified as leakage achieves high performance, outperforming human experts. These findings substantiate concerns that LJP performance may be exaggerated by linguistic artefacts. Yet this vulnerability is not fatal to the research agenda. Instead, post-hoc judgments might be treated as potentially contaminated texts, requiring active auditing. Retraining models after masking leakage features results in only a negligible reduction in Macro-F1. Hence, while models will opportunistically exploit shortcuts when available, they remain capable of extracting useful predictive signals when these artefacts are removed.
Theory of Mind (ToM) is a must-acquire skill for modern foundation model systems to operate effectively and safely in the real world. Recent works have explored honing ToM via post-training; however, we show that such progress is confounded by a pervasive "shortcut" issue: tasks can reach up to 99% accuracy by simply exploiting spurious causal correlations, leading to a false sense of ToM. Motivated by this, we first develop a framework to systematically examine ToM datasets for shortcuts and provide guidance for future development. We find that questions reducible to pure state tracking, such as "belief," are especially shortcut-prone compared to mind questions, such as "intention," where reasoning beyond tracking is required. Using four shortcut-free datasets across three ToM contexts, we then comprehensively study whether Reinforcement Fine-Tuning with verifiable rewards and explicit reasoning chains, called Thinking-RFT, elevates ToM beyond Supervised Fine-Tuning, or SFT. Our key findings are as follows. First, Thinking-RFT effectively improves ToM in all scenarios, with a 6% improvement over SFT, particularly in complex higher-order reasoning, with a 10% improvement over SFT, and multimodal cases, with a 7% improvement over SFT. It also generalizes notably better to unseen domains and higher-order queries while being more robust to counterfactuals. Second, ToM benefits specifically from the joint effect of reasoning and RL: Thinking-RFT outperforms Non-Thinking-RFT by 7% on average. Third, RFT works by learning to ground its reasoning on anchor cues, such as keywords and state changes, that correspond to causal factors. We believe our study is useful for developing effective and robust ToM post-training datasets and advancing critical ToM capabilities.
Fengyuan Liu, Yongliang Miao, Zirui He +3cs.LG cs.CL
Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framework for mitigating shortcut learning in reward model training. Unlike static shortcut heuristics, DynaCF measures shortcut sensitivity online during optimization by applying semantics-preserving counterfactual perturbations and tracking the resulting margin shifts and preference flips under the current model. Samples with higher shortcut sensitivity are dynamically downweighted in the Bradley-Terry objective, encouraging the model to rely less on superficial patterns and more on task-relevant preference signals. Extensive experiments show that DynaCF consistently improves robustness in preference modeling.