Epoch VI or the Radiologist

Inside our journey through an AI healthcare competition

Epoch VI, the new edition of the AI Dream Team that began this September, has emerged from our first competition! After six intense weeks of set up, planning and collaboration, the subteam of 4 (Aryaman Kusurkar, Laurens Ledeganck, Rein Viegers and Willem Dieleman) have submitted their final models. We are progressing now with important reflections and a deeper sense of what it means to work together under pressure against competitors at the top of the league. 

Choosing a Challenge That Matters

Societal Benefit

When choosing a challenge, our goal is not only to test technical skills, but work on something that matters. We look for projects that align with the UN Sustainable Development Goals and showcase how AI can genuinely improve lives. That’s what drew us to the Radiological Society of North America (RSNA) Intracranial Aneurysm Detection Challenge, a healthcare-focused competition centered on detecting brain aneurysms from medical scans. 

Aneurysms are a bulge or small balloon in the wall of a blood vessel. Intracranial (brain) aneurysms, which affect approximately 3% of the global population, often develop silently, showing no symptoms until they rupture and cause bleeding inside the brain at which point death is a common outcome. Even experienced radiologists can miss them on routine scans, with up to 50% being diagnosed only after they rupture. This is unfortunate considering that minimally-invasive procedures (avoiding open surgery) to treat intracranial aneurysms already exist. By training AI models to spot aneurysms across different imaging types and hospitals, the competition aimed to push AI toward tools that could one day help doctors catch these critical conditions early and save lives.

Role of AI in Radiology

Machine learning applications are available in multiple stages in neuroradiology (the study of the brain and nervous system using medical imaging): detection, diagnosis, treatment and post-treatment, as summarized in the diagram. At this stage, its role can be particularly significant in assistive work such as interpretation of medical images. In neuroradiology particularly, the Supervised ML and Deep Learning models are significant contributors to manage high dimensional data. Supervised ML involves pairing imaging data from X-rays, MRI or CTA scans with annotations from radiologists to train models to identify the pathology. Convolutional Neural Networks (CNNs), a form of Deep learning optimized for image processing, uses multiple layers in the network to effectively capture features (edges, textures, vessel structures) from high dimensional scans with millions of voxels (3D pixels). This is useful as the insights from imaging can be used to combine with clinical data to support doctors in developing effective treatment strategies. 

Machine learning application in neuroradiology

Competition Landscape

As this was the first competition we were participating in, we also wanted to select one that would be technically motivating. We understood this to be ideal considering:

  1. Unclear approach to the solution since the leaderboard showed the top competitors with scores around 0.88, 0.83, 0.81, 0.80, as opposed to scores that could be much higher and closer together within 0.01 ranges of each other.
  2. Very few of these top competitors, most remained below the 0.70 threshold which indicated there were good approaches that most people had not yet figured out.
  3. Application outside our area of expertise which would mean learning associated with managing medical data and files.
  4. Two month timeline, aligned with our internal plans.

The brain aneurysm challenge checked all the boxes!

Building the Strategy

Understanding the Data and Challenge

The sub team of 4 members:  focused on this specific competition began the week by waiting two hours to download 300GB of brain scans from 4300 patients, a larger than typical dataset for your average competition. Each patient’s data included one of four types of brain imaging scans, each showing different tissue or vessel details:

CTA (Computed Tomography Angiography) – Uses X-rays and a contrast dye to visualize blood vessels, key for aneurysm detection.
MRA (Magnetic Resonance Angiography) – A type of MRI that focuses on blood vessels using magnetic fields, without needing radiation.
T1 post-contrast MRI – shows general brain structure, highlights certain tissues or abnormalities after a contrast agent is injected.
T2-weighted MRI – Shows fluid and soft tissue details, helping to spot abnormalities like swelling or bleeding.

Each scan also came with demographic information (patient age and sex), and labels for 13 key blood vessels, indicating whether an aneurysm was present in each. The main target variable was a binary label: whether an aneurysm existed anywhere in the brain.

We also had to locate the abnormality, using segmentation. In medical AI, segmentation plays a central role. This means dividing an image into meaningful parts by labeling every pixel (or voxel in 3D scans) as a specific anatomical structure. This is crucial to locate abnormalities, beyond just detecting that they exist. Its accuracy can be measured using a dice score where 1 indicates perfect performance.

This setup, data exploration and generally researching the field we were in comprised the first two weeks.

The Early Breakthrough (or so we thought)

We spent the following weeks hooked on enhancing the existing public solutions because we saw it rank in the top 50. This was a CNN trained for image classification. One example of trying to optimize it was by compressing scans to the same uniform shape, but this shape could still include data from irrelevant areas in the scan as it did not actually recognize the region of interest – this would have required training a separate dedicated model. Eventually we realized progress was stalling, and decided to try new ideas.

This led to our first (supposed) success, which came not from the scans themselves, but from an interesting source – the metadata. This is the non-image information, such as scanner information, modality, rotation etc.. In week 3, we began training an Artificial Neural Network (ANN) using only metadata instead of the actual images – a Tabular Metadata model (TMM) which scored well, achieving 0.65.

After some reflection following this result, we knew we had to abandon the CNN approach. Considering how quickly it was able to reach a result, and performing very similarly to the TMM, we concluded it was also not really using the image data which is not a medically sound process.

A potential reason for this (that we did not have time to confirm) was that the model had learnt an accidental correlation: that if it’s a MR scan, there is likely an aneurysm. We knew this was not real medical insight but a result of how the dataset was distributed (possibly more aneurysm cases being in MR scans).

This was an important lesson: a high score doesn’t always mean the model has learned something meaningful. It was already week 4 of 6, and time to explore completely new ideas.

Pivoting to New Models

After intense brainstorming sessions within the team (and AI agents!), we moved on to image segmentation models. Half of the team worked on nnU-Net, a deep learning model specifically designed for medical image segmentation. It’s widely used in competitions like this because of its great performance and minimal setup, making it easy to adapt for our aneurysm detection challenge.

Using nnU-Net, we achieved the promising result of 0.71 Dice score for blood vessel segmentation, meaning it could reliably identify vessels. However, it was 0.3 Dice score for aneurysm detection, indicating it struggled with smaller or more subtle regions.

We also ran into practical issues. The model was too slow at inference; meaning it took too long to analyze new scans, and the competition constraint was just 17 seconds er patient to make a prediction. With only three days left before the deadline, we had to abandon this approach despite its potential.

Overview of the input data and model training process for Custom U-Net

Simultaneously, we were also building a custom U-Net, a simpler version of the nnU-Net but with more control – which was the approach we decided to then focus on. This was submitted in the end as it ended up being 1. the best performing model that did not make our computers crash – a requirement we set after exactly that happened, and 2. it ran in less than 12 hours on the submission server – a requirement of the competition.

Key Reflections

How We Fared (in context)

After some long nights in the final weeks we ended with a score of 0.68 out of 1. The top competitor scored 0.87, making us 611/1150. We were not surprised by this result, considering the fierce competition landscape. Top competitors include the following:

  1. So-called Grandmasters with years of experience and several international ML competition wins under their belt.
  2. Employees from large corporations like NVIDIA using competitions for their employee training, encouraged to place high on the leaderboard.
  3. For this specific competition, the runner up was even a group of PhD candidates based in China, specialized in the field of ML for healthcare.

Considering this, we are still extremely proud of our accomplishment since the strategy we developed in essentially one week aligned with the winning strategy (although we did not realize this at the time and had abandoned the approach). 

Real-World Challenges in Medical AI

There are challenges with implementation, including sufficient data volume and quantity. We were also able to comprehend within the scope of the competition considering how the TMM approach showed us that issues in our volume of data may have led to results that are not actually correct as they are not properly trained on real medical insight.

High effort is required to pre-process this data, which has a larger implication in hospitals with huge volumes of data. Research shows that an estimated 97% of unused hospital data can be used in AI applications for disease trajectory predictions and modifying treatment regimens.

Although we found that large improvements in the model performance itself are more dependent on the AI engineer rather than technical expertise from a medical profession, the collaboration is more important in reality. They can significantly contribute in providing integral context for the issue; be in considering real workflows, how can we actually make sure it is implemented efficiently and finally to support in assuring acceptance by the end users.

Furthermore, considering the sensitive nature of the application – the blackbox dilemma where AI models make predictions that we as humans cannot entirely understand. also becomes more significant – . It involves ethical considerations including data privacy, patient confidentiality, informed consent and prevention of misdiagnosis. It is also essential to ensure the training datasets used consider bias mitigation strategies to prevent unfair or inaccurate predictions for certain groups.

Moving Forward

In the end, the team walked away with lighter minds, as well as lighter servers (after clearing 8 TB from our machines). Overall, the competition was intense but fun! We learned a huge amount about the healthcare domain and got a glimpse of how much potential AI holds in transforming it. Along the way, we also developed new workflows, code, and intuition on what could (and could not) work. This experience is what we believe will make our next projects more efficient.

Beyond the technical side, it was also a great bonding experience, where we built the confidence to tackle large, unfamiliar challenges.

With these lessons in hand, we excitedly move on to the next competition!

Competition Sub Team: Rein, Laurens, Willem and Aryaman