Interleaving mitigates burst errors but introduces decoding delay and removes temporal error structure that a channel-aware decoder could exploit. We consider packet-level selection between a random linear code and the same code used with cross-codeword interleaving, over a channel with an unknown number of on/off interferers. The receiver uses Guessing Random Additive Noise Decoding (GRAND) with a replaceable noise model and feeds aggregate channel statistics back to a Bayesian estimator at the transmitter. Once the interference amplitudes and timing parameters are estimated, the receiver's noise model is replaced: it computes hidden-Markov-model posterior bit-flip probabilities and uses them to order GRAND queries. A discounted Thompson sampler selects between the two transmission modes using a goodput-minus-latency reward whose distribution is endogenously nonstationary: receiver adaptation, rather than channel change, alters the value of each mode. Across five simulation seeds, the interleaved mode is preferred before channel estimation converges. After the learned decoder is activated, the non-interleaved mode becomes preferable because it achieves lower block error rate without interleaving delay. In the reference configuration, the learned noise model reduces block error rate by approximately one order of magnitude relative to ORBGRAND. Using partial channel estimates before full convergence reduces pre-convergence block error rate by up to $4.5\times$. Adding model-predicted utilities as confidence-weighted pseudo-observations reduces post-transition selection of the inferior arm by approximately $65\%$. Under an idealized airtime conversion at a 100~MHz 5G~NR-like symbol rate, the learning transient corresponds to a few milliseconds of occupied symbol time.
Molecular communication (MC) suffers from severe diffusion memory because molecules released for one symbol may arrive during later symbol intervals. Neural sequence detectors, especially sliding bidirectional recurrent neural networks (SBRNNs), substantially outperform threshold detection in such channels. This raises a central question for MC channel coding: does a code whose superiority was established under threshold detection retain it when both coded and uncoded transmission are evaluated with neural detection? This letter answers this question for run-length-limited ISI-mitigation (RLIM) codes by proposing a decoder-aware training mask that removes the positions the RLIM decoder has a high probability of deterministically overwriting, steering compact-SBRNN capacity toward the information-bearing positions. The masked RLIM$_2$-SBRNN beats the best uncoded receiver (threshold or SBRNN) at 40 of 57 operating points; gains peak at 43$\times$ under favorable channels, while losses, confined to the most adverse, never exceed 2.7$\times$. Masking improves the unmasked RLIM$_2$-SBRNN in 56 of 57 matched comparisons. Finally, with storage counted equally in SBRNN weights and MLSE table entries, the masked RLIM$_2$-SBRNN is the more accurate receiver up to a few thousand stored values despite using no channel knowledge; channel-state-aware uncoded MLSE moves ahead only beyond tens of thousands.