Bulding a Community of Interest: discussing the IDEA4RC Ecosystem for Federated Rare Cancer Research with potential users

Cattura

Dissemination activities play a fundamental role in ensuring that the outcomes of European research projects extend beyond the lifetime of the projects themselves. By creating opportunities for dialogue with clinicians, researchers, healthcare organisations and technology experts, they transform research results into shared knowledge and encourage the adoption of innovative solutions capable of generating long-term scientific and societal impact.

It was within this framework that, on 28 May 2026, MultiMed Engineers, together with the Fondazione IRCCS Istituto Nazionale dei Tumori of Milan—the project’s coordinating institution—and the Azienda Ospedaliero-Universitaria Consorziale Policlinico di Bari, organised the workshop IDEA4RC: An Intelligent Data Ecosystem for Rare Cancers. Opportunities and Challenges when Deploying a Regulatory-Compliant System for Multi-Centre Secondary Use of Health Data.”

The event was hosted by the Unit of Maxillofacial Surgery, at the Policlinico di Bari.

Designed for a multidisciplinary audience—including medical oncologists, radiation oncologists, oncological surgeons, pathologists, clinical researchers, epidemiologists, data managers and ICT professionals—the workshop provided an opportunity to explore the IDEA4RC project within the concrete setting of a leading clinical centre specialising in the treatment and research of head and neck cancers and sarcomas. More importantly, it offered participants the opportunity to examine both the opportunities and the organisational challenges associated with joining the IDEA4RC ecosystem.

The event also represented an important dissemination milestone for the project. Alongside professionals from the organising institutions, participants included representatives from leading Italian healthcare and research organisations such as the IRCCS Istituto Tumori di Bari “Giovanni Paolo II” and the IRCCS Casa Sollievo della Sofferenza in San Giovanni Rotondo. Their participation reflected the growing interest in federated approaches to health data sharing and created an opportunity for direct dialogue between the project consortium and organisations that may contribute to the future expansion of the IDEA4RC ecosystem.

Throughout the afternoon, clinicians and researchers explored the opportunities offered by federated research infrastructures, discussed the perspectives opened by the European Health Data Space (EHDS) and reflected on the transformative potential of European collaboration in rare cancer research. The workshop therefore served not only as a forum for presenting the project’s achievements, but also as a platform for discussing how innovative digital infrastructures can help address one of the most challenging areas of contemporary oncology.

Rare cancers remain one of the greatest challenges for clinical research. Although each individual disease affects a relatively small number of patients, rare cancers collectively represent a significant public health issue. Their low incidence results in fragmented clinical knowledge, dispersed datasets and limited opportunities to build sufficiently large patient cohorts capable of supporting robust clinical studies or the development of reliable Artificial Intelligence models.

Funded by the European Union’s Horizon Europe Research and Innovation Programme under Grant Agreement No. 101057048, IDEA4RC addresses these challenges by developing a federated European ecosystem for the secure sharing and secondary use of health data in rare cancer research. Rather than centralising sensitive patient information, the project enables participating healthcare organisations to collaborate while maintaining full control over their own clinical data, thereby combining scientific collaboration with institutional data sovereignty and compliance with the evolving European regulatory framework.

The following topics were discussed during the Workshop.

Building a Federated Ecosystem for Collaborative Research

IDEA4RC addresses one of the fundamental challenges of rare cancer research: not simply the scarcity of data, but the ability to transform fragmented clinical information into secure, interoperable and comparable knowledge. The project achieves this through an integrated workflow that extends from locally generated hospital data to the controlled production of scientific evidence.

Unlike conventional centralised data-sharing approaches, the IDEA4RC ecosystem brings analytical capabilities to where the data already reside. Algorithms are executed locally within each participating institution, and only authorised outputs are returned for collaborative analyses. In other words, IDEA4RC moves algorithms to the data, rather than transferring data to a central repository.

This paradigm preserves institutional data sovereignty while enabling secure multicentre research. Healthcare organisations do not transfer their raw information assets; instead, they make them available for controlled and traceable use through a structured preparation process that encompasses structured and unstructured data sources, data quality assessment, standardised terminologies, a common data model and metadata harmonisation.

At the core of this architecture is the Local Capsule, the secure computational environment deployed within each participating institution. The Capsule protects both data and computation, allowing healthcare organisations to retain complete control over their information while making it available for federated research under clearly defined governance rules.

Rather than exporting clinical datasets, the Capsule stores harmonised data locally, exposes only the metadata required to support collaborative studies without revealing personal information, executes authorised algorithms within the hospital’s secure infrastructure, and enforces access policies, governance rules and restrictions on data extraction. This architecture enables collaboration without compromising privacy, security or institutional autonomy.

Integrating Structured and Unstructured Clinical Information

A distinctive feature of IDEA4RC is its ability to integrate highly heterogeneous clinical information.

Healthcare organisations routinely manage structured datasets, such as diagnoses, laboratory results, procedures and administrative records. However, an equally important body of clinical knowledge remains embedded within free-text documents, including electronic health record (EHR) notes, pathology reports, surgical reports and other narrative clinical documentation.

To address this challenge, IDEA4RC incorporates advanced Natural Language Processing (NLP) technologies capable of extracting clinically relevant information from narrative documents and transforming it into harmonised, computable data that can be incorporated into the Local Capsule alongside structured information.

An essential aspect of this process is the explicit representation of temporality.

In oncology, the clinical significance of an event depends not only on what happened but also on when it happened. Diagnoses, treatments, disease progression, recurrences and follow-up observations derive their clinical meaning from their chronological relationships.

For this reason, IDEA4RC has been designed to preserve the temporal dimension of clinical information throughout the extraction process. Clinical concepts identified within EHR notes and other narrative sources are transformed into computable data while maintaining their temporal relationships, enabling researchers to reconstruct clinically meaningful patient trajectories without losing traceability to the original documentation. Preserving temporality in this way is fundamental for generating research-ready datasets capable of supporting reliable clinical and epidemiological analyses.

Trusted Analytics and Reproducible Research

The workshop also highlighted the mechanisms that ensure transparency and trustworthiness throughout the analytical process.

Before execution, every statistical or Artificial Intelligence algorithm is catalogued, documented and formally approved. Each algorithm specifies its intended purpose, required inputs, expected outputs, operational limitations and conditions of use, while dedicated security and privacy controls minimise the risk of unintended information disclosure.

Equally important, comprehensive traceability and reproducibility mechanisms establish a direct link between workflows, algorithms and research outcomes, ensuring that analytical processes remain transparent, auditable and scientifically reproducible.

Finally IDEA4RC supports researchers in moving from a general scientific question to the identification of a computable and comparable patient cohort. To this end, the project develiped Raven—the Rare Cancer AI Virtual Exploration Navigator—the ecosystem’s research environment designed to enable the structured exploration of distributed clinical data while preserving interoperability, security and patient privacy.

From Deployment to Operational Experience

Drawing on the experience and lessons learned, accumulated throughout the project, seven key domains that every healthcare organisation should consider when preparing to join a federated research ecosystem like IDEA4RC have been discussed.

Governance as the Starting Point

The first lesson learned concerns governance.

Successful participation in a federated infrastructure begins long before any technical deployment takes place. The legal basis for participation must be established from the outset through the active involvement of the institution’s legal office, the Data Protection Officer (DPO) and the relevant clinical representatives. Together, these stakeholders define the conditions under which clinical data can be securely and legitimately used within the IDEA4RC framework.

Building on Secure Digital Foundations

A second key aspect concerns the local IT and cybersecurity environment.

The deployment strategy adopted by each institution depends largely on its existing technological infrastructure and security policies. Consequently, IT departments and cybersecurity specialists should be involved at an early stage to ensure that the deployment of the Local Capsule and the other IDEA4RC components is fully compatible with institutional requirements and available technical resources.

Data Harmonisation Before Data Sharing

One of the most important messages emerging from the pilot experience is that simply possessing clinical data is not sufficient.

To become useful for federated research, local datasets must be transformed into information that is compatible, semantically consistent and interoperable within the IDEA4RC Common Data Model.

This requires an iterative mapping process between local data sources and the common semantic model, carried out jointly by data managers and clinical experts. Only through continuous interaction between technical and clinical competencies can heterogeneous hospital information be successfully harmonised for collaborative research.

Data Quality as a Continuous Process

Another important lesson concerns data quality.

A dataset that is technically available is not necessarily suitable for federated analysis.

For this reason, participating institutions should establish a continuous quality improvement cycle that includes dataset preparation, automated quality assessment, analysis of feedback, correction of both data and semantic mappings, and repeated validation until an adequate level of quality has been achieved for collaborative analyses.

Within IDEA4RC, data quality is therefore not considered a final verification step but an ongoing process accompanying the entire lifecycle of research data.

Unlocking Clinical Knowledge Through NLP

The pilot implementations also highlighted the strategic importance of free-text clinical documentation.

Narrative reports frequently contain valuable clinical information that is unavailable in structured datasets. However, transforming this information into computable knowledge requires a carefully designed, authorised and clinically validated Natural Language Processing workflow.

Each participating institution must therefore identify which variables are available exclusively within narrative documentation and determine, together with clinical, technical and legal experts, whether these data can be extracted through NLP technologies or should remain outside the federated dataset.

Becoming an Active Node in Federated Research

Readiness for federated analysis is achieved progressively.

Institutions move from validated datasets and an operational Local Capsule to the execution of meaningful and interpretable collaborative analyses.

Becoming an active node within the IDEA4RC ecosystem means not only gaining the ability to query distributed datasets but also enabling authorised institutions to perform approved analyses on locally stored data through secure algorithm execution, while maintaining full institutional control over information assets.

Organisation Matters as Much as Technology

Perhaps the most significant conclusion presented during the workshop was that implementing IDEA4RC is not primarily a technological challenge.

Rather, it is an organisational process requiring multidisciplinary collaboration throughout the entire deployment pathway.

For this reason, every participating institution should establish a dedicated local implementation team capable of coordinating scientific, technical, legal and operational activities.

According to the experience gained during the project, such a team should include:

  • Scientific Lead
  • Data Manager
  • IT and Cybersecurity Lead
  • Legal Representative
  • Data Protection Officer (DPO)
  • Local Operational Coordinator

Once the deployment phase has been successfully completed, this multidisciplinary structure should evolve into a permanent operational coordination team capable of supporting the long-term sustainability of the local IDEA4RC infrastructure.

Beyond the Project: Building a Community of Interest for Federated Rare Cancer Research

The workshop concluded with an open and constructive discussion that confirmed the strong interest generated by the IDEA4RC ecosystem among healthcare professionals and researchers.

One of the aspects that attracted the greatest attention was the opportunity offered by the project’s federated architecture to enable secure data sharing while preserving institutional data sovereignty. By connecting a network of Local Capsules, participating Centres of Excellence can contribute their own rare cancer datasets—often consisting of only a limited number of cases—in exchange for access to a much broader pool of harmonised information made available by other participating institutions. This collaborative approach significantly increases the potential for conducting statistically robust studies and generating stronger scientific evidence for rare cancer research.

The discussion also highlighted the practical challenges that healthcare organisations may encounter when joining the IDEA4RC ecosystem.

Participants agreed that the most demanding aspect is often not the deployment of the technological infrastructure itself, but the preparation of clinical data originating from hospital information systems, particularly Electronic Health Records (EHRs). Data quality begins well before information enters the IDEA4RC ecosystem. It depends on how clinical information is generated, recorded, maintained and managed within each healthcare organisation.

At present, few institutions have professional roles specifically dedicated to these activities, while the personnel who currently perform them are frequently already engaged in demanding day-to-day operational responsibilities.

Importantly, this challenge is not unique to organisations approaching IDEA4RC for the first time.

The experience gained during the project has shown that several of the pilot Centres of Excellence encountered similar difficulties throughout the implementation process. In most cases, the barriers to joining a federated health data ecosystem are not related to a lack of scientific interest, but rather to practical constraints, including the availability of specialised data management personnel, adequate IT resources, an appropriate legal basis for data processing and the organisational effort required to coordinate multiple institutional stakeholders.

For organisations considering future participation, the workshop outlined a clear implementation pathway.

The first step is recognising the scientific and organisational value of joining the ecosystem. This must then be followed by a concrete institutional commitment to mobilise the necessary human, technical and organisational resources. Such a commitment requires coordinated action involving hospital management, legal departments, IT services and clinical researchers, all working towards a common objective.

Particular emphasis was placed on the data mapping process, which emerged as one of the most demanding phases of implementation.

Semantic mapping cannot be successfully carried out by technical specialists alone, nor exclusively by clinicians. It requires continuous collaboration between professionals who understand the clinical meaning of healthcare information and those responsible for translating that knowledge into interoperable digital models.

One of the most significant outcomes of four years of collaborative work within IDEA4RC has been precisely this process of mutual learning between clinical and technical professionals. Throughout the project, clinicians have gained a deeper understanding of the technological principles underpinning interoperable research infrastructures, while engineers and data specialists have become increasingly familiar with the clinical reasoning that gives meaning to healthcare data. This shared experience has fostered multidisciplinary teams capable of addressing complex research challenges through a common language and shared objectives.

It was with this spirit of collaboration that the workshop came to a close.

More than a dissemination event, the meeting demonstrated how technological innovation, scientific excellence and multidisciplinary cooperation can converge to support a common European vision for rare cancer research.

For MultiMed Engineers, this vision closely reflects our mission as a multidisciplinary team of engineers and medical experts collaborating in European research and innovation initiatives to develop advanced digital solutions for the biomedical and data-intensive domains. Our contribution to IDEA4RC is driven by the belief that meaningful innovation emerges from the integration of technological expertise, clinical knowledge and close collaboration with healthcare organisations.

The Workshop confirmed that the future of rare cancer research depends not only on innovative digital technologies, but also on the ability of clinicians, researchers, engineers and healthcare institutions to work together and form a Community of Interest within trusted, interoperable and sustainable European data ecosystems.

In this respect, IDEA4RC represents far more than a research project. It is helping to build the foundations of a collaborative European community capable of transforming fragmented health data into shared knowledge, accelerating scientific discovery and ultimately improving care for patients affected by rare cancers.