Direct deep learning analysis of three-dimensional automated breast ultrasound videos with reading mode optimization for breast cancer diagnosis
Article information
Abstract
Purpose
This study aimed to develop and evaluate a deep learning model that directly analyzes three-dimensional automated breast ultrasound videos (DL-3DABUV) to assist breast cancer diagnosis, and to examine the optimal reading mode for clinical implementation.
Methods
This retrospective study included 547 patients (285 benign, 262 malignant), who were randomly assigned to a training set (n=437) and a test set (n=110). The DL-3DABUV model, built using ResNet50 and multi-instance learning, was trained by directly analyzing videos without image selection or manual annotation. Six radiologists (three experienced and three novice) evaluated the test set under three modes: independent-reading (without DL-3DABUV), second-reading (without prior knowledge of DL-3DABUV results), and concurrent-reading (after viewing DL-3DABUV results). The diagnostic performance of DL-3DABUV, experienced radiologists, and novice radiologists was compared. Reading times across the three modes were also assessed.
Results
Compared to experienced radiologists in independent reading, DL-3DABUV showed no significant differences in area under the receiver operating characteristic curve (AUC) (0.82 vs. 0.83), sensitivity (82.1% vs. 81.6%), or specificity (81.5% vs. 88.3%) (all P>0.05). DL-3DABUV exhibited higher AUC and specificity than novice radiologists in independent-reading (0.82 vs. 0.68, P<0.001; 81.5% vs. 57.4%, P<0.001). However, novice performance reached parity with DL-3DABUV in both second-reading and concurrent-reading. No significant differences in diagnostic performance were observed between second-reading and concurrent-reading. Concurrent reading significantly reduced reading time by 33.6 seconds compared with second-reading (P<0.001).
Conclusion
DL-3DABUV achieves diagnostic performance comparable to experienced radiologists and enhances diagnostic accuracy for novices. Concurrent reading provides a more efficient workflow by reducing reading time while maintaining diagnostic performance.
Introduction
Breast cancer, the most common cancer among women, presents substantial challenges due to its high incidence and mortality [1]. Breast ultrasound provides high sensitivity, no radiation exposure, and cost-effectiveness, making it essential for diagnosis, especially in Asian women who commonly have dense breast tissue, a major risk factor [2,3]. It includes conventional two-dimensional handheld ultrasound (2DHHUS) and three-dimensional automated breast ultrasound (3DABUS). 2DHHUS requires manual probe manipulation, which contributes to operator dependency and poor reproducibility. In contrast, 3DABUS uses a wide-array transducer and robotic arm for automated scanning, rapidly acquiring standardized full-volume 3D reconstructions [4]. This generates sequential images in the transverse, sagittal, and coronal planes, reducing operator dependence, improving reproducibility, and providing important coronal views that support diagnosis [5]. However, the large number of images produced per patient can be time-consuming to interpret, potentially increasing false-positive findings and missed malignancies [6].
Recent advances in artificial intelligence (AI) may help address these challenges. AI-based computer-aided diagnosis systems have been proposed to improve both accuracy and efficiency in breast cancer diagnosis [7,8]. Deep learning, a central approach within AI that mimics the hierarchical structure of biological neural networks, autonomously extracts patterns from large labeled datasets through iterative layer training. This process builds predictive models without requiring manual feature engineering [9]. Convolutional neural networks (CNNs) demonstrate strong feature extraction and learning capabilities with minimal human intervention [10]. They can directly learn features from input data, making them highly effective for classifying 3DABUS images [11–15]. However, most existing models were trained on selected static images or required expert annotation. This raises concerns about image selection, since key diagnostic features may appear only in certain tumor views, and suboptimal selection can lead to misinterpretation [16]. In addition, using only one or a few images may fail to capture the complete tumor characteristics, resulting in incomplete analysis and loss of the spatiotemporal context inherent in ultrasound. The substantial manual annotation burden is also time-consuming and labor-intensive, limiting model scalability.
To address these limitations, it was hypothesized that a deep learning model that directly analyzes 3DABUS videos can achieve high diagnostic performance for breast cancer. This approach uses the full imaging dataset, avoids labor-intensive manual annotations, and aligns with clinical workflows. A recent study using contrast-enhanced ultrasound videos reported 86.3% diagnostic accuracy, demonstrating the potential of video-based deep learning for diagnostic support [17]. AI systems can also assist radiologists as a second reader (used after initial interpretation, potentially increasing reading time [18]) or as a concurrent reader (used at the start, which may reduce specificity due to lower vigilance [19]). However, the optimal strategy for incorporating these reading modes into clinical workflows for AI-assisted 3DABUS interpretation remains unclear.
Therefore, this study aims to develop and evaluate a deep learning model that directly analyzes 3DABUS videos (DL-3DABUV) to assist breast cancer diagnosis and to investigate the optimal reading mode for clinical integration.
Materials and Methods
Compliance with Ethical Standards
This retrospective study was approved by the Institutional Review Board of the authors’affiliated hospital (No. 2022K104), and the requirement for written informed consent was waived.
Patient Population
A retrospective review was conducted of data from patients who underwent 3DABUS between June 2012 and December 2022. All malignant cases and a subset of benign cases were pathologically confirmed via biopsy or surgery, while the remaining benign cases were confirmed by ≥2 years of follow-up at the hospital with no interval changes. Inclusion criteria were as follows: (1) presence of 3DABUS-identified breast masses; (2) benign diagnoses requiring ≥2 years of follow-up at the hospital; and (3) malignant diagnoses requiring pathological confirmation by biopsy or surgery. Exclusion criteria were as follows: (1) simple cysts (defined as circumscribed, round or oval, anechoic masses with posterior enhancement [20]); (2) nonmass lesions (which are difficult to identify as discrete masses on ultrasound [21]); (3) prior breast surgery or implants before 3DABUS; (4) pre-existing radiotherapy, chemotherapy, or targeted therapy; (5) suboptimal 3DABUS images with artifacts; and (6) incomplete clinical information. Ultimately, 547 patients with complete 3DABUS video data were included: 262 with malignant and 285 with benign masses. Among the 285 benign cases, 193 (67.7%) were pathologically confirmed and 92 (32.3%) were confirmed by ≥2 years of follow-up. Fig. 1 illustrates the research process.
3DABUS Videos
The ACUSON S2000 Automated Breast Volume Scanner System (Siemens Medical Solutions, Mountain View, CA, USA) equipped with a 14L5BV 5–14 MHz linear array transducer was used. Patients were positioned in a supine or lateral decubitus position with arms naturally raised to ensure full breast exposure. Scanning depth and gain were adjusted based on breast size to allow clear visualization from the nipple to the chest wall. Anteroposterior and lateral scans were routinely performed for full breast coverage, and supplementary views (such as medial, superior, or inferior) were obtained when needed. The acquired volumetric data were transferred to a dedicated 3DABUS workstation for reconstruction. Using a standardized protocol, the software produced uniform 3D datasets for each acquisition, consistently generating 318 transverse slices at a 0.5 mm interval regardless of breast size or scan depth. Coronal and sagittal series were also reconstructed at uniform 0.5 mm intervals.
The acquisition of 3DABUS videos involved collecting breast volumetric data using the 3DABUS device, transferring DICOM sequences to the workstation for standardized 3D reconstruction in all three planes, and generating dynamic playback sequence videos for each plane. Fig. 2 illustrates this process, and representative videos are provided in Video clips 1 and 2. The 3DABUS videos were used both for training the AI model (DL-3DABUV) and for simulating radiologist interpretation in the reader study, during which radiologists navigated the dynamic sequences to evaluate lesion characteristics, mirroring real-world diagnostic workflows. This study utilized 3DABUS videos consisting of continuous transverse and coronal sequences. Transverse–coronal plane selection was based on three factors: (1) the unique value of coronal reconstruction in 3DABUS, which provides diagnostic information unattainable with 2DHHUS and enhances early detection in dense breasts [4,5]; (2) the use of transverse–coronal dual-plane analysis as the primary diagnostic protocol, with sagittal views as supplements; and (3) historical limitations in systematic video storage that restricted available planes to these two views.
DL-3DABUV Development
DL-3DABUV development included data preprocessing and model architecture design (Fig. 3). 3DABUS videos from 547 patients were randomly divided into training (437 patients) and test (110 patients) sets at a 4:1 ratio. The training set consisted of 231 benign (52.9%) and 206 malignant (47.1%) cases, while the test set included 54 benign (49.1%) and 56 malignant (50.9%) cases. During preprocessing, all frames were extracted from transverse and coronal planes of the 3DABUS videos. Patient information and other interfering elements were removed by edge cropping. Images were resized to standard dimensions and normalized for consistency. Data augmentation techniques (rotation, flipping, cropping) were applied to improve model robustness. Two input strategies were constructed: full-frame sequences (all video frames) and tumor-frame sequences (frames containing masses).
DL-3DABUV architecture.
Ti-n, Ti-m, Ti, Ti+m, ..., Ti+n indicate input frames at different time points. 3DABUS, three-dimensional automated breast ultrasound; 3DABUV, three-dimensional automated breast ultrasound videos; DL-3DABUV, deep learning for three-dimensional automated breast ultrasound videos.
For the model architecture, an ImageNet-pretrained ResNet50 backbone was used to leverage transfer learning for faster training and improved performance. During training, the original fully connected layer was replaced with linear layers and ReLU activation to output benign or malignant mass probabilities. Transverse and coronal frames were processed separately to generate single-frame predictions. A multiple instance learning (MIL) framework was incorporated, treating single-frame predictions as instances and entire videos as bags. This allowed aggregation of intra-bag information to fully utilize sequential frame features. Specifically, the highest malignant probability predicted for any single frame within a video was taken as the malignant probability for the entire video. Predicted probabilities from the coronal and transverse videos of each case were then weighted and averaged to produce a final predicted probability. A threshold determined by maximizing Youden’s Index was applied to categorize lesions as benign or malignant. Integrating ResNet50 with MIL enhanced feature learning and improved robustness to label noise. This combination utilizes ResNet50’s strong feature extraction capabilities and MIL’s ability to model inter-instance relationships within bags [22,23], improving classification performance while reducing annotation burden.
Grad-CAM++ is an advanced method for visualizing and interpreting CNN decision-making. It generates attention maps by computing a weighted sum of the gradients of the output with respect to the activations of a convolutional layer, enabling more precise localization of features influencing model predictions, particularly when multiple regions show similar activation. The resulting heatmap highlights areas most influential to the model’s decisions. In this study, Grad-CAM++ was used to generate class activation heatmaps for test-set frames based on the ResNet50 backbone, illustrating the contribution of different image regions to predicted outputs. Supplementary Fig. 1 presents representative heatmaps, showing that DL-3DABUV’s activation was predominantly concentrated within tumor regions. Heatmap intensity (with red indicating high activation) correlated with gradient weights, reflecting stronger network activation and greater diagnostic importance. These findings confirm that DL-3DABUV consistently focuses on diagnostically relevant areas, thereby improving interpretability.
Implementation was performed using the PyTorch framework on an Ubuntu 20.04 system equipped with an Intel(R) Xeon(R) Platinum 8269CY CPU @ 2.50 GHz and an NVIDIA GeForce RTX 3090 24 GB GPU. The classification model used an initial learning rate of 1×10⁻³ and a batch size of 80. Network training was performed on the training set, with parameters optimized using the test set.
Reader Study
Six radiologists participated in the reader study and were categorized into an experienced group (radiologists 1, 2, and 3) and a novice group (radiologists 4, 5, and 6). Their experience profiles were as follows: Radiologist 1 had 7 years of ultrasound experience, including 5 years specializing in breast ultrasound, and performed and interpreted approximately 500 breast ultrasound examinations and 150 3DABUS examinations annually. Radiologists 2 and 3 each had 5 years of ultrasound experience with 3 years specializing in breast ultrasound, and each performed and interpreted approximately 600 breast ultrasound examinations and 200 3DABUS examinations annually. Radiologists 4 and 5 each had 2 years of ultrasound experience, including 1 year specializing in breast ultrasound, and each performed and interpreted approximately 400 breast ultrasound examinations and 100 3DABUS examinations annually. Radiologist 6 had 1 year of ultrasound experience following dedicated breast ultrasound training and performed and interpreted approximately 300 breast ultrasound examinations and 60 3DABUS examinations annually. Three reading modes were employed. In the independent-reading mode, radiologists reviewed 3DABUS videos independently to determine the final diagnosis. In the second-reading mode, radiologists initially reviewed 3DABUS videos without DL-3DABUV results, after which the DL-3DABUV results were provided and incorporated into the final diagnosis. In the concurrent-reading mode, radiologists first accessed the DL-3DABUV results and subsequently reviewed the 3DABUS videos before determining the diagnosis. Each reading mode was separated by an interval of more than 3 months to minimize memory effects [24]. Before the study, all radiologists completed training on the reading modes using 30 3DABUS video cases (excluded from the study cohort), with 10 cases allocated to each mode. Several measures were implemented to ensure objectivity. First, radiologists were permitted to view only the 3DABUS videos and their corresponding case numbers. Second, the 3DABUS videos were randomly renumbered before each reading session. Third, each radiologist viewed the 3DABUS videos in a unique randomized order.
According to the 2013 Breast Imaging Reporting and Data System (BI-RADS) established by the American College of Radiology, masses categorized as BI-RADS category 3 were defined as benign, whereas those categorized as BI-RADS category 4A or higher were defined as malignant [20]. In this reader study, a benign case was defined as a 3DABUS video in which all masses were assigned BI-RADS category 3. A malignant case was defined as a 3DABUS video containing at least one mass categorized as BI-RADS 4A or higher, even if additional category 3 masses were present. Additionally, the time each radiologist required to complete the diagnosis in each reading mode was recorded. Reading time was defined as the interval from the moment a radiologist accessed a video to the time the diagnosis was finalized. Radiologists were blinded to the timing mechanism throughout the process.
Statistical Analysis
Diagnostic performance was evaluated using the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, and accuracy. Pairwise comparisons of diagnostic performance were conducted between DL-3DABUV and both experienced and novice radiologists across the three reading modes. AUC values were compared using the DeLong test. Comparisons of sensitivity and specificity were conducted using the McNemar test. Differences in management decision changes between the second-reading and concurrent-reading modes were evaluated using the Mann-Whitney U test. Reading times across different reading modes were compared using the independent-samples t-test. All statistical analyses were performed using SPSS version 22.0 (IBM Corp., Armonk, NY, USA) and statistical significance was defined as P<0.05.
Results
Patient Characteristics
This study enrolled 547 female patients (age range, 22 to 83 years; mean±standard deviation, 50.8±14.3 years; median, 50 years). A total of 730 breast lesions were identified. Patients were categorized into a benign group (285 patients, 440 lesions) and a malignant group (262 patients, 290 lesions). Within the benign group, 168 patients had single lesions and 117 had multiple lesions. In the malignant group, 240 patients had single lesions and 22 had multiple lesions. Lesion size ranged from 7 to 46 mm (median, 13 mm; mean±standard deviation, 14.9±7.0 mm).
DL-3DABUV Diagnostic Performance Evaluation
The diagnostic performance of DL-3DABUV, experienced radiologists across the three reading modes, and novice radiologists across the three reading modes is summarized in Table 1. As shown in Table 1 and Supplementary Table 1, DL-3DABUV demonstrated diagnostic performance comparable to that of experienced radiologists in independent-reading mode, with no significant differences in AUC (0.82 vs. 0.83, P=0.230), sensitivity (82.1% vs. 81.6%, P>0.99), or specificity (81.5% vs. 88.3%, P=0.126). Compared with novice radiologists in independent-reading mode, DL-3DABUV showed significantly higher AUC (0.82 vs. 0.68, P<0.001) and specificity (81.5% vs. 57.4%, P<0.001), while maintaining comparable sensitivity (82.1% vs. 79.2%, P=0.590). When assisted by DL-3DABUV, radiologist performance varied according to experience. Experienced radiologists achieved significantly better diagnostic performance in both second-reading and concurrent-reading modes compared with DL-3DABUV alone (all P<0.01 for AUC, sensitivity, and specificity; Table 1 and Supplementary Table 1). In contrast, novice radiologists in these modes achieved AUC (second-reading: 0.84 vs. 0.82, P=0.196; concurrent-reading: 0.85 vs. 0.82, P=0.125), sensitivity (second-reading: 87.5% vs. 82.1%, P=0.211; concurrent-reading: 88.1% vs. 82.1%, P=0.154), and specificity (second-reading: 82.7% vs. 81.5%, P=0.856; concurrent-reading: 83.3% vs. 81.5%, P=0.720) that were comparable to DL-3DABUV (Fig. 4).
ROC curves for DL-3DABUV, experienced radiologists (three modes), and novice radiologists (three modes).
ROC, receiver operating characteristic curve; AUC, area under the ROC curve; DL-3DABUV, deep learning for three-dimensional automated breast ultrasound; ERs, experienced radiologists; NRs, novice radiologists.
Impact of Reading Mode on Diagnostic Performance
According to Table 1 and Supplementary Table 1, both second-reading and concurrent-reading modes significantly improved diagnostic performance compared with independent-reading for both experienced and novice radiologists. Experienced radiologists achieved significantly higher AUC (second-reading: 0.90 vs. 0.83, P<0.001; concurrent-reading: 0.90 vs. 0.83, P<0.001), sensitivity (second-reading: 90.5% vs. 81.6%, P<0.001; concurrent-reading: 91.7% vs. 81.6%, P<0.001), and specificity (second-reading: 92.0% vs. 88.3%, P=0.031; concurrent-reading: 92.0% vs. 88.3%, P=0.016) than in independent-reading. Similarly, novice radiologists also showed significantly improved AUC (second-reading: 0.84 vs. 0.68, P<0.001; concurrent-reading: 0.85 vs. 0.68, P<0.001), sensitivity (second-reading: 87.5% vs. 79.2%, P<0.001; concurrent-reading: 88.1% vs. 79.2%, P<0.001), and specificity (second-reading: 82.7% vs. 57.4%, P<0.001; concurrent-reading: 83.3% vs. 57.4%, P<0.001) compared with independent-reading. However, there were no significant differences between second-reading and concurrent-reading in AUC (experienced radiologists: 0.90 vs. 0.90, P=0.500; novice radiologists: 0.84 vs. 0.85, P>0.99), sensitivity (experienced radiologists: 90.5% vs. 91.7%, P>0.99; novice radiologists: 87.5% vs. 88.1%, P>0.99), or specificity (experienced radiologists: 92.0% vs. 92.0%, P=0.156; novice radiologists: 82.7% vs. 83.3%, P=0.157). Across all three reading modes, experienced radiologists demonstrated significantly higher AUC (independent-reading: 0.83 vs. 0.68, P<0.001; second-reading: 0.90 vs. 0.84, P=0.008; concurrent-reading: 0.90 vs. 0.85, P=0.006) and specificity (independent-reading: 88.3% vs. 57.4%, P<0.001; second-reading: 92.0% vs. 82.7%, P=0.014; concurrent-reading: 92.0% vs. 83.3%, P=0.020) than novice radiologists, while maintaining comparable sensitivity (independent-reading: 81.6% vs. 79.2%, P=0.644; second-reading: 90.5% vs. 87.5%, P=0.424; concurrent-reading: 91.7% vs. 88.1%, P=0.286). For novices, both reading modes with DL-3DABUV performed comparably to the independent-reading performance of experienced radiologists (Fig. 4).
Management Decision Changes
Radiologists’ BI-RADS category adjustments with DL-3DABUV assistance are summarized in Table 2. For experienced radiologists, all downgrades involved changes from BI-RADS category 4A to category 3, with six cases in each mode. The final diagnoses of these cases were benign, indicating that six unnecessary biopsies (1.8%) were correctly avoided. Upgrades included changes from BI-RADS category 3 to category 4A in 14 cases during second-reading and 13 cases during concurrent-reading. One case was upgraded from category 3 to category 4B in second-reading, whereas four such upgrades occurred during concurrent-reading. The final diagnoses of these upgraded cases were malignant, indicating that 15 additional cancers (4.5%) were detected in second-reading and 17 (5.1%) in concurrent-reading. For novice radiologists, downgrades primarily involved changes from BI-RADS category 4A to category 3, with 39 cases downgraded in second-reading and 41 cases in concurrent-reading. Second-reading also included 2 downgrades from category 4B to category 3, while concurrent-reading had 1 similar downgrade. All downgraded cases were ultimately benign, indicating that 41 unnecessary biopsies (12.4%) were avoided in second-reading and 42 (12.7%) in concurrent-reading. Upgrades for novice radiologists included 10 cases changed from BI-RADS category 3 to category 4A in second-reading and seven cases in concurrent-reading. Additionally, four upgrades from category 3 to category 4B occurred in second-reading, while eight upgrades occurred in concurrent-reading. These upgraded cases were malignant, indicating that 14 additional cancers (4.2%) were detected in second-reading and 15 in concurrent-reading (4.5%). Fig. 5 illustrates representative cases of BI-RADS category adjustments with DL-3DABUV assistance.
BI-RADS adjustments with DL-3DABUV assistance.
A, B. Case 1 (invasive carcinoma, 56-year-old woman): transverse (A) and coronal (B) images are shown. Arrows in (A) and (B) indicate the mass. Experienced radiologists reported the following assessments in independent-reading: Radiologist 1, BI-RADS category 4C; Radiologist 2, BI-RADS category 3; Radiologist 3, BI-RADS category 4B. Novice radiologists reported the following assessments in independent-reading: Radiologist 4, BI-RADS category 4B; Radiologist 5, BI-RADS category 4B; Radiologist 6, BI-RADS category 3. After DL-3DABUV review, Radiologist 2 and Radiologist 6 upgraded from BI-RADS category 3 to 4A. C, D. Case 2 (fibroadenoma, 47-year-old woman): Transverse (C) and coronal (D) images are shown. Arrows in (C) and (D) indicate the mass. Experienced radiologists reported the following assessments in independent reading: Radiologist 1, BI-RADS category 3; Radiologist 2, BI-RADS category 3; Radiologist 3, BI-RADS category 4A. Novice radiologists reported the following assessments in independent-reading: Radiologist 4, BI-RADS category 4A; Radiologist 5, BI-RADS category 4A; Radiologist 6, BI-RADS category 4B. After DL-3DABUV review, Radiologist 3, 4 and 5 downgraded from BI-RADS category 4A to 3, and Radiologist 6 downgraded from BI-RADS category 4B to 3. BI-RADS, Breast Imaging Reporting and Data System; DL-3DABUV, deep learning for three-dimensional automated breast ultrasound.
Reading Times
For all radiologists, mean reading times ± standard deviation were 137.2±23.2 seconds (range, 100 to 188 seconds) in independent-reading mode, 160.6±23.9 seconds (range, 120 to 204 seconds) in second-reading mode, and 127.1±18.5 seconds (range, 96 to 160 seconds) in concurrent-reading mode. As shown in Supplementary Fig. 2, reading times were longest in second-reading mode and shortest in concurrent-reading mode across experienced radiologists, novice radiologists, and the combined cohort. Compared with both independent-reading and second-reading, concurrent-reading resulted in significantly shorter reading times for all radiologists, with mean differences of 13.5 seconds (95% confidence interval [CI], 12.7 to 14.2; P<0.001) and 33.6 seconds (95% CI, 32.5 to 34.6; P<0.001), respectively (Table 3).
Discussion
This work explored a novel deep learning strategy for direct analysis of 3DABUS videos and developed an AI system (DL-3DABUV) that achieved diagnostic performance comparable to that of experienced radiologists in breast cancer diagnosis. By processing continuous video sequences, DL-3DABUV captured comprehensive spatiotemporal information without requiring extensive manual annotation. DL-3DABUV showed diagnostic performance comparable to that of experienced radiologists (AUC, 0.82 vs. 0.83; P=0.230) and significantly improved the performance of novice radiologists. Clinically, the concurrent-reading mode proved more efficient than the second-reading mode, providing similar diagnostic performance while significantly reducing reading time.
A major strength of this study lies in the systematic validation of DL-3DABUV’s diagnostic efficacy and its collaborative value when integrated with radiologist interpretation. The reader study demonstrated that DL-3DABUV performed comparably to experienced radiologists (sensitivity, 82.1% vs. 81.6%; specificity, 81.5% vs. 88.3%) and outperformed novices (sensitivity, 82.1% vs. 79.2%; specificity, 81.5% vs. 57.4%), findings consistent with previous reports [11,14,15,25–28]. DL-3DABUV significantly improved diagnostic performance for both experienced and novice groups, which aligns with prior studies of AI-assisted diagnostic systems [18,29]. This study further evaluated AI integration strategies (second-reading vs. concurrent-reading) within clinical 3DABUS workflows by analyzing reading time. Both strategies demonstrated comparable diagnostic performance (all P>0.1). Concurrent reading was the fastest across all groups, whereas second-reading took the longest, consistent with other AI systems [7,18]. Offering equivalent diagnostic efficacy while saving 33.6 seconds per case, concurrent reading represents an optimized workflow strategy. Notably, although experienced radiologists outperformed DL-3DABUV once assisted by the system, with the mean AUC significantly improving from 0.82 to 0.90, novice radiologists achieved performance comparable to DL-3DABUV alone (Table 1, Supplementary Table 1). These findings highlight a complex interaction between radiologist expertise and AI assistance. Although DL-3DABUV significantly improved novice performance to a level similar to the AI itself (AUC, 0.84–0.85 vs. 0.82; P>0.05), this raises the question of whether radiologist interpretation remains necessary for less experienced readers when AI support is readily available. DL-3DABUV may function as a preliminary screening tool, especially in regions lacking expert radiologists [30], by triaging clearly benign cases and flagging suspicious ones for further evaluation, thereby improving workflow efficiency. However, radiologists remain indispensable. First, the significant performance gains among experienced radiologists indicate that optimal diagnostic accuracy is achieved through human–AI synergy. Second, for novices, AI-assisted reading also serves an educational function, allowing their diagnostic judgments to be calibrated against AI output and potentially accelerating their long-term skill development. The reduced reading time in concurrent reading may reflect this educational benefit and an increase in diagnostic confidence. Finally, radiologists perform tasks beyond binary classification, including lesion characterization, integration of clinical history, biopsy recommendations, and communication of findings—responsibilities that remain firmly within human expertise [30]. Thus, DL-3DABUV provides dual utility: in resource-limited settings, it can function as a standalone screening aid that mitigates expertise gaps, whereas in well-equipped centers it functions as a decision-support tool that enhances radiologists’ performance and improves overall diagnostic accuracy.
Another strength of this study is the detailed evaluation of BI-RADS category adjustments with DL-3DABUV assistance. Experienced radiologists mainly made upgrades (4.5% and 5.1%), suggesting that DL-3DABUV helps identify subtle malignant features and supports early detection. In contrast, novices predominantly made downgrades (12.4% and 12.7%). Their specificity increased from 57.4% (independent-reading) to 82.7% (second-reading) and 83.3% (concurrent-reading). The lower specificity of novices in independent reading may reflect a defensive diagnostic strategy, arising from limited experience, in which ambiguous lesions are overclassified to avoid missed cancers [31]. Breast ultrasound has historically been criticized for high false-positive rates [32]. The substantial number of downgrades with DL-3DABUV assistance suggests that DL-3DABUV mitigates over-classification, thereby potentially reducing unnecessary biopsies. Sensitivity was comparable between experienced (81.6%) and novice (79.2%) radiologists in independent-reading and remained high in DL-3DABUV–assisted modes (experienced, 90.5% and 91.7%; novice, 87.5% and 88.1%). This likely reflects 3DABUS’s high lesion detectability enabled by automated scanning and multiplanar reconstruction [2,5]. Nevertheless, novice specificity remained consistently lower than that of experienced radiologists across reading modes (independent-reading, 57.4% vs. 88.3%; assisted, 82.7% vs. 92.0% and 83.3% vs. 92.0%), demonstrating that clinical experience is still critical for diagnostic accuracy, even when AI support is available (Table 1, Supplementary Table 1).
Importantly, this study offers methodological advancements for deep learning in medical image analysis. Unlike previous breast AI systems that relied on manually annotated and selected static images with image-level or pixel-level labels [11–15,25–28], this system directly analyzes 3DABUS videos using sequence-level labels, eliminating per-frame annotation. This approach provides three major advantages: reduced annotation burden, complete utilization of the imaging dataset, and improved alignment with clinical workflows. Such efficiency is essential for the development of clinically scalable AI systems, particularly given the need for large-scale validation across diverse clinical contexts where detailed annotation is impractical. This method therefore establishes a scalable and data-efficient paradigm for future AI development.
Notably, radiologists were limited to 3DABUS videos and patient age, whereas clinical practice typically incorporates prior imaging and medical histories [2]. This standardized setting controlled key variables by evaluating DL-3DABUV’s impact within a baseline environment consisting only of imaging and age, thereby avoiding bias from historical clinical data. It also reflects conditions encountered in initial screening or in resource-limited settings where prior records may be unavailable. Although reduced clinical context limits generalizability to fully comprehensive clinical environments, DL-3DABUV nonetheless demonstrated substantial diagnostic utility under these constraints, supporting its potential role in screening workflows and in settings with limited clinical information [30]. Future research should develop AI capabilities for extracting features from historical examinations and integrating multimodal longitudinal data.
Several limitations warrant consideration. First, the retrospective single-center design may limit generalizability and introduce recruitment and selection biases; multicenter prospective validation is needed. Second, the 3DABUS videos lacked sagittal-plane images, and future studies incorporating all three orthogonal planes may further improve diagnostic performance. Finally, the present analysis focused on patient-level rather than lesion-level evaluation. Clinically, the presence of any malignant lesion mandates appropriate intervention, and this approach aligns with current patient management strategies that prioritize the highest-suspicion finding. However, lesion-level analysis could enable more precise localization and targeted biopsy guidance, and such approaches should be explored in future work.
In conclusion, this study demonstrated the feasibility of direct deep learning analysis of 3DABUS videos. The model achieved diagnostic performance comparable to experienced radiologists, and collaboration between radiologists and DL-3DABUV significantly enhanced diagnostic performance, particularly among novices. This direct deep learning analysis strategy—leveraging full-sequence utilization and eliminating the need for manual annotation—offers valuable insights for future AI development. Additionally, the concurrent-reading mode optimized workflow efficiency compared with the second-reading mode.
Notes
Author Contributions
Conceptualization: Guo Y, Wang C, Zhang Q, Chen L. Data acquisition: Guo Y, Wang C, Liu Y, Pang Y, Ge R, Li W, Chen L. Data analysis or interpretation: Guo Y, Wang C, Liu L, Zhang Q, Chen L. Drafting of the manuscript: Guo Y, Wang C. Critical revision of the manuscript: Liu Y, Pang Y, Ge R, Li W, Liu L, Zhang Q, Chen L. Approval of the final version of the manuscript: all authors.
Conflict of Interest
No potential conflict of interest relevant to this article was reported.
Acknowledgments
This work was supported by grants for a 2022 Science and Technology Project from the Technology Innovation Action Plan (Project No. 22Y11911600, Chen L) and a 2020 National Natural Science Foundation of China Project (Project No. 62071285, Zhang Q). The authors gratefully acknowledge the six radiologists (Xiaoli Peng, Yumeng Xing, Jiamei Jin, Yuhan Liu, Chunhong Ding, and Linwen Xu) for their valuable participation in the reader study. Special thanks are extended to Chunyan Zhong and Haier Wang for their meticulous assistance in time-recording procedures during the experimental phase.
Supplementary Material
Pairwise comparison of diagnostic performance among DL-3DABUV, ER, and NR across three reading modes (https://doi.org/10.14366/usg.25096).
Representative Grad-CAM++ heatmaps on 3DABUS images (https://doi.org/10.14366/usg.25096).
Reading time distribution by radiologist and mode (https://doi.org/10.14366/usg.25096).
Video clip 1.
Video of the transverse view from three-dimensional automated breast ultrasound of a 56-year-old woman with left breast invasive carcinoma (https://doi.org/10.14366/usg.25096.v1).
Video clip 2.
Video of the coronal view from three-dimensional automated breast ultrasound of a 56-year-old woman with left breast invasive carcinoma (https://doi.org/10.14366/usg.25096.v2).
References
Article information Continued
Notes
Key points
A deep learning model that directly analyzes three-dimensional automated breast ultrasound videos without image selection or manual annotation was developed. The model demonstrated diagnostic performance comparable to that of experienced radiologists and significantly improved the diagnostic accuracy of novices. Concurrent-reading mode reduced reading time while maintaining diagnostic performance comparable to second-reading mode, indicating a more efficient workflow strategy.
