Robust and Quality-of-Life-Aware Treatment Protocols in NSCLC using Deep Reinforcement Learning
Journal:
bioRxiv
Published Date:
Sep 4, 2026
Abstract
Under current systemic treatment of metastatic cancer, a drug is frequently prescribed at maximum tolerable dose (MTD) until either unacceptable toxicity or progression. Unfortunately, in many patients this treatment strategy leads to the development of treatment resistance. Evolutionary therapy approaches aim to forestall or delay treatment resistance in cancer by exploiting eco-evolutionary interactions. A well-known implementation is the adaptive therapy protocol of Zhang et al., in which tumour burden thresholds are used to guide strategic treatment holidays. Deep reinforcement learning (DRL) has recently been used to optimise these approaches. However, research combining DRL with evolutionary therapy approaches has so far focused on time to progression (TTP) as a performance metric, and has not quantified safety in terms of robustness to delayed treatment restart or included patient preferences regarding quality of life (QoL) in treatment design. In our study, we use a DRL agent informed by a mathematical two-population tumour growth model to design treatment schedules for patients with non-small cell lung cancer (NSCLC). The agent is trained on a virtual patient cohort using parameters previously fitted to data from patients with NSCLC treated with erlotinib. Beyond TTP, we focus on improving robustness to delayed treatment restart and on how individual preferences and values impact QoL experienced during treatment. We compare TTP, robustness and QoL under the DRL policy, the adaptive therapy protocol of Zhang et al., and MTD. We introduce a robustness metric ``margin-to-failure'' (MTF), and compare quality-adjusted-survival (QAS) across different patient preference profiles. Finally, we explore reward shaping to assess how QoL preferences can be incorporated into DRL-based treatment design. To evaluate our results, we consider different decision intervals, defined as the time between dosing adjustments. The DRL policy achieved greater median TTP, MTF, and QAS across all treatment decision intervals compared to the other two protocols. As decision intervals increased, TTP under DRL declined gradually towards that achieved under MTD. In contrast, the Zhang et al. protocol performed inconsistently and could result in premature progression. Additionally, a population-level policy trained on a cohort of virtual patients produced an interpretable treatment rule that extended TTP for most previously unseen patients and indicated that treatment should resume at a lower tumour burden when monitoring is less frequent. These findings show that DRL can balance the benefit of preserving drug-sensitive cells to suppress resistance against the risk of unsafe tumour regrowth. Reward shaping further showed how treatment strategies could be adjusted to reflect different patient preferences. Together, these results provide a biologically informed approach for designing robust and patient-centred evolutionary therapies in fast-growing cancers such as NSCLC.