どうも、Beyond the Pixelです。胎児心臓超音波検査は、先天性心疾患の早期発見に不可欠な非侵襲的かつリアルタイムな画像診断法です。しかし、診断に用いる標準的な心臓断面(Standard Plane)の取得は、高度な専門知識と熟練した技術を要するだけでなく、非常に時間のかかる作業です。この手動によるプロセスは、超音波専門家の疲労や、複数の解剖学的ビュー間で非常に似通った空間的特徴から微妙な局所的画像変化を識別するという困難さから、その自動化が長年の課題となっていました。例えば、
Figure 1. Fig. 1. Inter- and intra-annotator variability analysis for the task of standard plane detection in a transverse heart sweep. It reveals that the average agreement between the annotators is 66%, underscoring the challenge of this task.解説: この図は、横断スウィープにおける標準断面検出タスクにおける超音波専門家間の評価者内および評価者間変動性分析を示しています。10個のスキャンIDについて、評価者内(Intra)と評価者間(Inter)の一致率がカッパスコアで表示されており、平均一致率が66%であることを強調し、このタスクの難しさを示しています。
Figure 2. Fig. 2. A plot of standard plane detection model accuracy versus the number of model parameters for the four models compared in the paper (see Section 5.2). Circle size indicates accuracy, with larger circles representing higher parameters.解説: この図は、本論文で比較された4つのモデル(提案モデル、ConvNeXt、DeiT、VOLO)における標準断面検出モデルの精度とモデルパラメータ数の関係を示しています。円のサイズは精度を示し、円が大きいほど高精度であることを表します。提案モデルは、他のモデルと比較して少ないパラメータ数で高い精度を達成していることを示唆しています。
Figure 3. Fig. 3. Overview of the framework for classifying five fetal heart standard planes (SITUS, 4CHV, LVOT, 3VV, and 3VTV) from t-sweep or freehand video frames.解説: この図は、胎児心臓の5つの標準断面(SITUS、4CHV、LVOT、3VV、3VTV)をt-sweepまたはフリーハンドビデオフレームから分類するためのフレームワークの概要です。入力ビデオフレーム(224x224x3ピクセル)がハーモニックブロックとscSEブロックで構成される層を通過し、平均プーリングと全結合層を経て最終的な分類出力が得られるプロセスを示しています。
HarmonicEchoNetモデルの開発と評価には、PULSE(Perception Ultrasound by Learning Sonographic Experience)とCAIFE(Clinical Artificial Intelligence in Fetal Echocardiography)という2つの私的臨床研究から得られた4つのデータセットが使用されました。これらのデータセットは、英国のNHSからの実際の対象者データで構成されており、プライバシー保護のため公開はされていませんが、確立された臨床取得プロトコルに従って収集されました。
Table 1. Table 1 Dataset sizes, reported as the number of frames, for the four normal (healthy) fetal heart datasets used in this study, excluding PULSE-Freehand para-standard. Refer to the main text for an explanation of each dataset. N is the number of unique videos and F is number of unique fetuses. Heart view/ Datasets Classes CAIFE t-sweep (N = 90, F = 19) standard解説: この表は、本研究で使用された4つの正常な(健康な)胎児心臓データセット(CAIFE t-sweep標準、CAIFE t-sweepパラ標準、CAIFEフリーハンド標準、PULSEフリーハンド標準)のデータサイズをフレーム数で報告しています。各データセットについて、SITUS、4CHV、LVOT、3VV、3VTVの各心臓ビューのトレーニング、バリデーション、テストセットにおけるフレーム数が示されています。
Figure 4. Fig. 4. Illustration of the two different data acquisition protocols. For the t-sweep (upper row), the sonographer records a 10-s video clip containing five fetal heart standard planes in a fixed sequential order (SITUS, 4CHV, LVOT, 3VV, and 3VTV). For the freehand case (lower row), the sonographer records short clips for e ach standard plane independently and in arbitrary order (opportunistic).解説: この図は、2つの異なるデータ取得プロトコルを図解しています。上段のt-sweep(横断スウィープ)では、超音波専門家が胎児心臓の5つの標準断面(SITUS、4CHV、LVOT、3VV、3VTV)を含む約10秒間のビデオクリップを固定された順序で記録します。下段のフリーハンドスキャンでは、超音波専門家が各標準断面の短いクリップを独立して任意の順序で記録します。
Figure 5. Fig. 5. The transverse sweep (t-sweep). (I) The situs view of the upper abdomen is visualized first. (II) The four-chamber view is obtained through an axial scanning plane across the fetal chest by moving and tilting the transducer in a cephalad direction. Further, cephalad movement of the transducer from the four-chamber view towards the fetal head gives the outflow-tract and great-vessel views sequentially: (III) left ventricular outflow-tract view; (IV) right ventricular outflow-tract view and the three- vessel view variants; and (V) three-vessel-and-trachea view.解説: この図は、トランスバーススウィープ(t-sweep)によって取得される胎児心臓の標準断面を視覚的に示しています。(I)腹部上部のSITUSビューから始まり、(II)胎児胸部を横断する軸方向スキャン面で四腔断面が取得され、プローブを頭側に動かすことで(III)左室流出路断面、(IV)右室流出路断面と三血管断面、そして(V)三血管気管断面が順に得られる様子が示されています。
Table 2. Table 2 Dataset sizes, reported as the number of frames, N is the number of unique videos for the PULSE pretraining dataset. Heart view/Classes PULSE pretraining dataset (N = 357, F = 357) Train Val Test SITUS 34 465 2366 12 558 4CHV 14 833 818 5094 LVOT 23 023 1456 7553 3VV 13 650 819 5187 3VTV 19 656 1274 7007解説: この表は、PULSE事前学習データセットのデータサイズをフレーム数で報告しています。SITUS、4CHV、LVOT、3VV、3VTVの各心臓ビューについて、トレーニング、バリデーション、テストセットにおけるフレーム数が示されており、事前学習のために使用されたデータの規模を表しています。
Figure 6. Fig. 6. Confusion matrices using ConvNeXt (Liu et al., 2022), DeiT (Touvron et al., 2021), VOLO (Yuan et al., 2022), and the proposed model for classifying the fetal heart standard plane in US video frames.解説: この図は、ConvNeXt、DeiT、VOLO、および提案モデル(HarmonicEchoNet)を用いた胎児心臓標準断面の分類における混同行列を示しています。各行列は、PULSE-Freehand標準データセットにおける各心臓ビュー(SITUS、4CHV、LVOT、3VV、3VTV)の実際のクラスと予測されたクラスの割合を表しており、HarmonicEchoNetが他のモデルと比較して高い正解率を示していることがわかります。
Table 4. Table 4 Computational complexity: Comparison of HarmonicEchoNet, ConvNeXT, DeiT and VOLO models. The best results are shown in bold. These models were built using the CAIFE t-sweep para-standard dataset. Architecture Parameters (M) FLOPS (G) MACs (G) Inference Time (S) ConvNeXt 196.24 135.39 67.65 0.0101 DeiT 303.35 238.34 119.13 0.0137 VOLO 293.92 270.39 135.12 0.0277 HarmonicEchoNet (Ours) 19.9 18.16 9.06 0.0039解説: この表は、HarmonicEchoNet、ConvNeXt、DeiT、VOLOモデルの計算複雑度を比較しています。パラメータ数(M)、FLOPS(G)、MACs(G)、推論時間(S)の各指標が示されており、HarmonicEchoNetモデルがパラメータ数、FLOPs、MACs、推論時間の全てにおいて他のモデルよりも優れた効率性を持っていることが示されています。
Table 5. Table 5 Performance comparison of CNN- and HC-based models to select the model with the best architecture. These models were trined on the CAIFE t-sweep para-standard dataset. The best values are in bold. Metrics are accuracy (AC), precision (PR), recall (RE), and F1-score (F1) reported as percentages. Architecture AC PR RE F1 Parameters (M) CNN (baseline) 86.19 86.22 83.85 84.83 18.51 CNN + scSE 86.39 86.34 84.67 85.25 19.9 HC (baseline) 90.31 88.82 87.47 88.05 18.51 HC+hscSE (proposed) 92.99 91.99 90.88 91.30 19.9解説: この表は、HarmonicEchoNetの最も優れたアーキテクチャを選択するためのCNNベースおよびHCベースモデルの性能比較を示しています。CAIFE t-sweepパラ標準データセットでトレーニングされたこれらのモデルについて、精度(AC)、適合率(PR)、再現率(RE)、F1スコア、およびパラメータ数(M)が比較されており、HC+hscSE(提案モデル)が最高の性能を達成していることが示されています。
Table 7. Table 7 Evaluating the effect of data augmentation on HarmonicEchoNet model performance. Metrics are accuracy (AC), precision (PR), recall (RE), and F1-score (F1) reported as percentages. Data Aug. AC PR RE F1 No 91.32 90.89 89.10 89.94 Yes 92.99 91.99 90.88 91.30
に示すように、適用しない場合に比べて精度が約1.1%から1.67%向上しました。損失関数としては、
Table 8. Table 8 Evaluating the effect of the loss function on the performance of a HarmonicEchoNet model trained on the CAIFE t-sweep para-standard dataset. Metrics are accuracy (AC), precision (PR), recall (RE) and F1-score (F1) reported as percentages. Loss function AC PR RE F1 CE 92.62 91.46 90.33 91.09 Focal 92.99 91.99 90.88 91.30
Table 9. Table 9 Evaluating the effect of pre-training on HarmonicEchoNet model performance across different datasets with and without pre-training on the PULSE freehand para-standard dataset. Model Dataset Pre-trained AC PR RE F1 CAIFE t-sweep standard No 69.24 67.60 66.71 64.47 Yes 77.05 79.78 73.21 71.51 CAIFE t-sweep para-standard No 68.45 73.31 67.78 69.24 Yes 76.42 81.85 80.48 80.32 CAIFE freehand standard No 78.88 79.20 76.98 77.58 Yes 89.11 89.02 88.04 88.42
Figure 7. Fig. 7. HarmonicEchoNet misclassification examples: from left to right image human annotation (GT) is LVOT, 3VV, and 3VTV, and the predictions (PD) are 4CHV, 3VTV, and 3VV, respectively.
Figure 8
Figure 8. Fig. 8. Visualization of the different intermediate layers for HarmonicEchoNet (upper row) and baseline CNN models without hscSE blocks (lower row). From these visualizations, observe that the harmonic convolutional layers appear to learn certain patterns through a neighborhood weighted sum that allows the layer to capture 4CHV image detail well. In comparison, the convolutional layers for the baseline do not seem to learn them as well.
Figure 9
Figure 9. Fig. 9. Activation map visualization. Each column shows (a) an input fetal heart standard plane, (b) the corresponding activation maps from the outputs of the ConvNeXt model (Liu et al., 2022), and (c) HarmonicEchoNet model. Observe that the HarmonicEchoNet model activation maps qualitatively align better with the region of interest associated with that heart view.
Figure 10
Figure 10. Fig. 10. T-SNE feature visualization of the fetal heart standard views classification by ConvNeXt (Liu et al., 2022), DeiT (Touvron et al., 2021), VOLO (Yuan et al., 2022), and HarmonicEchoNet models.
Leave a Reply