Got to News overview News

From Single Cells to Actionable Insights: How WP7 Powers Data Integration in PERSIST-SEQ

News News

Why do some cancer cells survive therapy—and how do they eventually give rise to drug resistance? Answering this question requires not only cutting-edge experimental models, but also a robust analytical framework capable of integrating complex single-cell data across technologies, models, and timepoints. 

Within PERSIST-SEQ, Work Package 7 (WP7) provides exactly this foundation. By developing standardized, reproducible bioinformatic pipelines and by collaborating closely with experimental teams, WP7 transforms raw sequencing data into biological insights that help decode the mechanisms underlying drug tolerance and resistance. 

A shared analytical backbone for the consortium 

PERSIST-SEQ generates large volumes of single-cell sequencing data across a wide range of experimental models, including cell lines, co-cultures, xenograft models, and patient-derived samples. To make meaningful comparisons across this diversity, WP7 has established a unified and reproducible data-analysis pipeline that is used consistently throughout the consortium. 

Once single cells have been sequenced, every dataset follows the same structured workflow: 

  • First, sequencing data undergoes rigorous quality control (QC) to ensure reliability. Raw reads are then aligned to a standardized reference genome, producing gene expression tables. These tables are combined with detailed experimental metadata into a harmonized dataset, following an agreed-upon file and folder structure.
  • The resulting datasets are stored securely in AWS S3 cloud storage and an HPC Archive, ensuring both accessibility and long-term preservation. 

From exploration to interpretation 

To enable rapid and consistent exploration of newly generated data, WP7 uses the web-based analysis platform Galaxy. Galaxy automatically picks up newly uploaded datasets and performs an initial set of standardized exploratory analyses. This step provides bioinformaticians with a first overview of the data, helping to identify trends, technical issues, or promising biological signals. 

From there, WP7 moves into more in-depth analyses, combining computational expertise with biological interpretation. Results and next steps are shared transparently through monthly work package meetings and dedicated Slack channels, allowing continuous feedback between computational and experimental teams. 

In practice, WP7 oversees the full computational lifecycle of PERSIST-SEQ data: 

  • Quality control and raw read alignment
  • Dataset assembly and standardization
  • Exploratory analysis using Galaxy pipelines
  • In-depth biological interpretation
  • Secure archiving and documentation 

Developing new methods to meet new challenges 

Some of the questions tackled within PERSIST-SEQ require analytical solutions that go beyond existing tools. A clear example is the challenge of analysing xenograft models, where human tumour cells or cells from cell cultures are implanted into a mouse model. 

To address this, the consortium developed a new computational approach using the 10X Flex platform, applying both human and mouse probe sets simultaneously. By assessing the proportion and specificity of probes detected in each cell, WP7 can reliably distinguish true human tumour cells from mouse-derived cells. This method enables more accurate downstream analyses of tumour–microenvironment interactions and has been described in a recent preprint

Making diverse data comparable 

One of WP7’s most important roles is ensuring that data generated across different labs, technologies and model systems can be meaningfully integrated. To minimize technical variability and batch effects, PERSIST-SEQ samples are processed centrally at Single Cell Discoveries, following standardized protocols and sequenced on the same Illumina machines. 

On the computational side, reads are aligned using identical software versions and reference genomes, while all processing scripts are version-controlled via GitHub. The combination of standardized wet-lab workflows, harmonized data formats, and shared analysis tools allows the consortium to focus on biological variation with as minimal as possible technical noise

The complexity of multimodal integration 

Despite extensive standardization, integrating single-cell data across various models, cancer types and timepoints remains challenging. Batch effects can arise not only from technical sources but also from genuine biological differences between samples. This is particularly pronounced in drug-persistence experiments, where pre-treatment and post-treatment cells necessarily come from different biological states. 

Another challenge is statistical power. Highly sensitive drug-response models often produce very few persister cells, limiting sample sizes and requiring careful interpretation. In addition, each experimental system brings its own trade-offs: cell lines offer control but lack tumour microenvironment interactions, while patient samples capture clinical reality but introduce additional genetic heterogeneity. 

Finally, time-resolved analysis adds further complexity. Capturing both rapid adaptive responses and slower transitions into stable resistance states requires careful experimental design, and where timepoints are limited, computational interpolation becomes essential. 

From data patterns to therapeutic hypotheses 

Despite these challenges, WP7’s analytical framework enables the consortium to uncover biologically meaningful insights. By mapping differentially expressed genes and proteins onto known pathways, WP7 identifies cellular processes that become dysregulated during resistance development. 

Graph-based algorithms help pinpoint signalling hubs or bottlenecks that may represent therapeutic vulnerabilities, while unsupervised clustering reveals distinct resistance subtypes that could require different treatment strategies. Trajectory and pseudotime analyses provide a temporal perspective, helping to reconstruct how resistance emerges and to identify early molecular changes that may precede clinically detectable relapse. 

A continuous dialogue between computation and experiments 

WP7 works hand-in-hand with experimental work packages throughout the project. Experimental teams share data in standardized formats, while WP7 performs initial analyses and flags potential issues early. Joint meetings ensure that computational results are interpreted within the correct biological context and that future experiments are informed by analytical insights. 

One concrete outcome of this collaboration comes from cell–cell communication analyses in co-culture models. WP7 identified specific ligand–receptor interactions that may support the survival of drug-persistent cells. These candidate pathways are now being tested experimentally using targeted perturbation and blocking or rescue approaches, closing the loop between computation and biology. 

Looking ahead: lasting impact beyond the project 

WP7’s work is designed with sustainability in mind. All raw data and analytical tools will be released alongside publications, enabling reuse by the wider scientific community. At the same time, the project delivers both new computational approaches and deeper biological insights into persister cell behaviour. 

Together, these outputs contribute to a more detailed understanding of therapy resistance, forming an essential step towards developing treatments that not only respond to resistance, but prevent it from emerging in the first place. 

Share this page…