SEARCH
検索詳細
松永 卓明大学院医学系研究科 医科学専攻助教
研究活動情報
■ 論文- OBJECTIVE: This study aims to evaluate whether large language models (LLMs) can accurately predict the urgency and severity of radiology reports. MATERIALS AND METHODS: Based on the recommendations of the Academy of Royal Colleges, we defined radiology reports that include unexpected findings of high urgency or severity as "high-priority (HP) radiology reports." Overall, 1906 radiology reports were used as the training set, and 176 radiology reports were used as the test set, with a balanced ratio of HP to non-HP radiology reports (1:1) in both sets. Four types of LLMs (Llama2 7B, Llama3 8B, Llama3 Elyza 8B, and Llama 3.1 8B) were fine-tuned using four different input settings: (1) findings only, (2) findings + referring department, (3) findings + referring department + clinical diagnosis before examination, and (4) findings + referring department + clinical diagnosis before examination + details of examination request. The fine-tuned LLMs predicted whether each radiology report was HP or not. RESULTS: Among the four LLMs, Llama3 Elyza 8B, with inputs comprising findings and the referring department, demonstrated the best performance, achieving PRAUC = 0.962, ROCAUC = 0.968, accuracy = 0.915, sensitivity/recall = 0.932, specificity = 0.898, and F1 = 0.916. Adding a clinical diagnosis before the examination and details of examination requests did not necessarily lead to performance improvement. CONCLUSION: The fine-tuned LLMs accurately predicted HP radiology reports, suggesting their potential utility in supporting communication regarding radiology reports with high urgency or severity. KEY POINTS: Question This study aims to evaluate whether large language models (LLMs) can accurately predict the high-priority (HP) radiology reports. Findings The fine-tuned best LLM accurately HP radiology reports, achieving PRAUC of 0.962 and ROCAUC of 0.968. Clinical relevance This study demonstrates that fine-tuned LLMs can accurately identify HP radiology reports, potentially improving timely clinical decision-making and enhancing patient safety through faster communication of critical findings.2025年12月, European radiology, 英語, 国際誌研究論文(学術雑誌)
- With the increasing use of postmortem imaging, deep learning (DL)-based automated analysis may assist in the detection of intracranial hemorrhages. However, limited postmortem data complicate model training. This study aims to assess the accuracy of DL models in detecting intracranial hemorrhages in postmortem head computed tomography (CT) scans using transfer learning. A total of 75,000 labeled head CT images from the Radiological Society of North America Intracranial Hemorrhage Detection Challenge serve as the training data for the 15 DL models. Each model is fine-tuned via transfer learning. A total of 134 postmortem cases with hemorrhage status confirmed by autopsy serve as the external test set. Model performance is evaluated using the area under the receiver operating characteristic curve (AUC), sensitivity, specificity, training time, inference time, and number of parameters. Spearman’s rank correlation coefficients are calculated for these metrics. DenseNet201 achieves the highest AUC (0.907), with the AUCs of the 15 models ranging from 0.862 to 0.907. A longer inference time moderately correlates with higher AUC (Spearman’s ρ = 0.586, p = 0.022), whereas the number of parameters is not positively correlated with performance (ρ = −0.472, p = 0.076). The sensitivity and specificity are 0.828 and 0.871, respectively. Transfer learning using a large non-postmortem dataset enables accurate intracranial hemorrhage detection using postmortem CT, potentially reducing the autopsy workload. The results demonstrate that models with fewer parameters often perform comparably to more complex models, emphasizing the need to balance accuracy with computational efficiency.MDPI AG, 2025年09月, Applied Sciences, 15(19) (19), 10513 - 10513研究論文(学術雑誌)
- BACKGROUND/OBJECTIVES: This study aimed to investigate the accuracy of Tumor, Node, Metastasis (TNM) classification based on radiology reports using GPT3.5-turbo (GPT3.5) and the utility of multilingual large language models (LLMs) in both Japanese and English. METHODS: Utilizing GPT3.5, we developed a system to automatically generate TNM classifications from chest computed tomography reports for lung cancer and evaluate its performance. We statistically analyzed the impact of providing full or partial TNM definitions in both languages using a generalized linear mixed model. RESULTS: The highest accuracy was attained with full TNM definitions and radiology reports in English (M = 94%, N = 80%, T = 47%, and TNM combined = 36%). Providing definitions for each of the T, N, and M factors statistically improved their respective accuracies (T: odds ratio [OR] = 2.35, p < 0.001; N: OR = 1.94, p < 0.01; M: OR = 2.50, p < 0.001). Japanese reports exhibited decreased N and M accuracies (N accuracy: OR = 0.74 and M accuracy: OR = 0.21). CONCLUSIONS: This study underscores the potential of multilingual LLMs for automatic TNM classification in radiology reports. Even without additional model training, performance improvements were evident with the provided TNM definitions, indicating LLMs' relevance in radiology contexts.2024年10月, Cancers, 16(21) (21), 英語, 国際誌研究論文(学術雑誌)
- Elsevier BV, 2024年06月, Informatics in Medicine Unlocked, 46, 101465 - 101465研究論文(学術雑誌)
- RATIONALE AND OBJECTIVES: Pericardial fat (PF)-the thoracic visceral fat surrounding the heart-promotes the development of coronary artery disease by inducing inflammation of the coronary arteries. To evaluate PF, we generated pericardial fat count images (PFCIs) from chest radiographs (CXRs) using a dedicated deep-learning model. MATERIALS AND METHODS: We reviewed data of 269 consecutive patients who underwent coronary computed tomography (CT). We excluded patients with metal implants, pleural effusion, history of thoracic surgery, or malignancy. Thus, the data of 191 patients were used. We generated PFCIs from the projection of three-dimensional CT images, wherein fat accumulation was represented by a high pixel value. Three different deep-learning models, including CycleGAN were combined in the proposed method to generate PFCIs from CXRs. A single CycleGAN-based model was used to generate PFCIs from CXRs for comparison with the proposed method. To evaluate the image quality of the generated PFCIs, structural similarity index measure (SSIM), mean squared error (MSE), and mean absolute error (MAE) of (i) the PFCI generated using the proposed method and (ii) the PFCI generated using the single model were compared. RESULTS: The mean SSIM, MSE, and MAE were 8.56 × 10-1, 1.28 × 10-2, and 3.57 × 10-2, respectively, for the proposed model, and 7.62 × 10-1, 1.98 × 10-2, and 5.04 × 10-2, respectively, for the single CycleGAN-based model. CONCLUSION: PFCIs generated from CXRs with the proposed model showed better performance than those generated with the single model. The evaluation of PF without CT may be possible using the proposed method.2024年03月, Academic radiology, 31(3) (3), 822 - 829, 英語, 国際誌研究論文(学術雑誌)
- IEEE, 2024年, 48th IEEE Annual Computers, Software, and Applications Conference(COMPSAC), 1952 - 1954研究論文(国際会議プロシーディングス)
- BACKGROUND AND PURPOSE: Mean pulmonary artery pressure (mPAP) is a key index for chronic thromboembolic pulmonary hypertension (CTEPH). Using machine learning, we attempted to construct an accurate prediction model for mPAP in patients with CTEPH. METHODS: A total of 136 patients diagnosed with CTEPH were included, for whom mPAP was measured. The following patient data were used as explanatory variables in the model: basic patient information (age and sex), blood tests (brain natriuretic peptide (BNP)), echocardiography (tricuspid valve pressure gradient (TRPG)), and chest radiography (cardiothoracic ratio (CTR), right second arc ratio, and presence of avascular area). Seven machine learning methods including linear regression were used for the multivariable prediction models. Additionally, prediction models were constructed using the AutoML software. Among the 136 patients, 2/3 and 1/3 were used as training and validation sets, respectively. The average of R squared was obtained from 10 different data splittings of the training and validation sets. RESULTS: The optimal machine learning model was linear regression (averaged R squared, 0.360). The optimal combination of explanatory variables with linear regression was age, BNP level, TRPG level, and CTR (averaged R squared, 0.388). The R squared of the optimal multivariable linear regression model was higher than that of the univariable linear regression model with only TRPG. CONCLUSION: We constructed a more accurate prediction model for mPAP in patients with CTEPH than a model of TRPG only. The prediction performance of our model was improved by selecting the optimal machine learning method and combination of explanatory variables.2024年, PloS one, 19(4) (4), e0300716, 英語, 国際誌研究論文(学術雑誌)
- 2024年, Nihon Hoshasen Gijutsu Gakkai zasshi, 80(6) (6), 673 - 678, 日本語, 国内誌研究論文(学術雑誌)
- 2023年12月, Proceedings of the 17th NTCIR Conference on Evaluation of Information Access Technologies, NTCIR, 155 - 162[査読有り]研究論文(国際会議プロシーディングス)
- To evaluate the diagnostic performance of our deep learning (DL) model of COVID-19 and investigate whether the diagnostic performance of radiologists was improved by referring to our model. Our datasets contained chest X-rays (CXRs) for the following three categories: normal (NORMAL), non-COVID-19 pneumonia (PNEUMONIA), and COVID-19 pneumonia (COVID). We used two public datasets and private dataset collected from eight hospitals for the development and external validation of our DL model (26,393 CXRs). Eight radiologists performed two reading sessions: one session was performed with reference to CXRs only, and the other was performed with reference to both CXRs and the results of the DL model. The evaluation metrics for the reading session were accuracy, sensitivity, specificity, and area under the curve (AUC). The accuracy of our DL model was 0.733, and that of the eight radiologists without DL was 0.696 ± 0.031. There was a significant difference in AUC between the radiologists with and without DL for COVID versus NORMAL or PNEUMONIA (p = 0.0038). Our DL model alone showed better diagnostic performance than that of most radiologists. In addition, our model significantly improved the diagnostic performance of radiologists for COVID versus NORMAL or PNEUMONIA.2023年10月, Scientific reports, 13(1) (1), 17533 - 17533, 英語, 国際誌研究論文(学術雑誌)
- PURPOSE: The purpose of this study is to compare two libraries dedicated to the Markov chain Monte Carlo method: pystan and numpyro. In the comparison, we mainly focused on the agreement of estimated latent parameters and the performance of sampling using the Markov chain Monte Carlo method in Bayesian item response theory (IRT). MATERIALS AND METHODS: Bayesian 1PL-IRT and 2PL-IRT were implemented with pystan and numpyro. Then, the Bayesian 1PL-IRT and 2PL-IRT were applied to two types of medical data obtained from a published article. The same prior distributions of latent parameters were used in both pystan and numpyro. Estimation results of latent parameters of 1PL-IRT and 2PL-IRT were compared between pystan and numpyro. Additionally, the computational cost of the Markov chain Monte Carlo method was compared between the two libraries. To evaluate the computational cost of IRT models, simulation data were generated from the medical data and numpyro. RESULTS: For all the combinations of IRT types (1PL-IRT or 2PL-IRT) and medical data types, the mean and standard deviation of the estimated latent parameters were in good agreement between pystan and numpyro. In most cases, the sampling time using the Markov chain Monte Carlo method was shorter in numpyro than that in pystan. When the large-sized simulation data were used, numpyro with a graphics processing unit was useful for reducing the sampling time. CONCLUSION: Numpyro and pystan were useful for applying the Bayesian 1PL-IRT and 2PL-IRT. Our results show that the two libraries yielded similar estimation result and that regarding to sampling time, the fastest libraries differed based on the dataset size.2023年, PeerJ. Computer science, 9, e1620, 英語, 国際誌研究論文(学術雑誌)
- Liver fibrosis is one of the common complications of transient myeloproliferative disorder (TMD) in Down syndrome (DS), but the exact molecular pathogenesis is largely unknown. We herein report a neonate of DS with liver fibrosis associated with TMD, in which we performed the serial profibrogenic cytokines analyses. We found the active monocyte chemoattractant protein-1 expression in the affected liver tissue and also found that both serum and urinary monocyte chemoattractant protein-1 concentrations are noninvasive biomarkers of liver fibrosis. We also showed a prospective of the future anticytokine therapy with herbal medicine for the liver fibrosis associated with TMD in DS.2017年07月, Journal of pediatric hematology/oncology, 39(5) (5), e285-e289, 英語, 国際誌研究論文(学術雑誌)
- 2026年, 臨床放射線, 71(3) (3)肝胆膵の画像診断基本とアップデート 肝胆膵の画像診断におけるAIの現状
- 京都 : 日本放射線技術学会, 2024年06月, 日本放射線技術学会雑誌 = Japanese journal of radiological technology, 80(6) (6), 673 - 678, 日本語放射線技術学研究におけるPythonの活用術 応用編(11)胸部単純X線写真の診断レポートの自動作成
- 2024年, 日本医学放射線学会秋季臨床大会抄録集, 60th死亡時画像診断(Ai)の人工知能(AI)
- (公社)日本医学放射線学会, 2022年03月, 日本医学放射線学会学術集会抄録集, 81回, S232 - S232, 英語深層学習を用いた肺結節の三次元CT画像の生成(Generation of Three-Dimensional CT Images of Lung Nodules using Deep Learning)
- 空置術後の内腸骨動脈瘤に対して透視下上臀動脈直接穿刺併用にて塞栓し得た1例症例は80歳代男性。主訴は突然の腹痛であった。造影CTで腹部大動脈瘤、両側総腸骨動脈瘤、両側内腸骨動脈瘤を認め、径73mm大の左総腸骨動脈瘤破裂と診断された。緊急でY-graft置換術を施行され、Y-graftの脚は左外腸骨動脈と右総腸骨動脈にそれぞれ吻合された。この際27mm大の左内腸骨動脈瘤を認めていたが、緊急手術であったため根部の結紮のみが行われた。7ヵ月後に右総腸骨動脈瘤と右内腸骨動脈瘤に対するステントグラフト内挿術および右内腸骨動脈のコイル塞栓術が施行された。グラフと置換術後3年で左内腸骨動脈の瘤径が拡大し、42mmに増大したため塞栓術が依頼された。空置術後で順行性のアクセスが不可能となった内腸骨動脈瘤に対して、側副路経由の経動脈アプローチに加えて仰臥位のまま透視下に上臀動脈を直接穿刺し、双方向アプローチにて内腸骨動脈瘤の塞栓術を行い成功した。(一社)日本インターベンショナルラジオロジー学会, 2022年03月, 日本インターベンショナルラジオロジー学会雑誌, 36(2) (2), 137 - 141, 日本語
- 2022年, 日本インターベンショナルラジオロジー学会雑誌(Web), 36(2) (2)空置術後の内腸骨動脈瘤に対して透視下上殿動脈直接穿刺併用にて塞栓し得た1例
- (一社)日本医療情報学会, 2021年11月, 医療情報学連合大会論文集, 41回, 1122 - 1124, 日本語大学病院における遺伝的アルゴリズムを用いた当直予定表作成システムの開発
- (一社)日本インターベンショナルラジオロジー学会, 2021年04月, 日本インターベンショナルラジオロジー学会雑誌, 36(Suppl.) (Suppl.), 234 - 235, 日本語腸管虚血を合併した上腸間膜動脈閉塞症に対して開腹下でIVRを施行した3例
- (一社)日本インターベンショナルラジオロジー学会, 2021年04月, 日本インターベンショナルラジオロジー学会雑誌, 36(Suppl.) (Suppl.), 254 - 255, 日本語当院における肝細胞癌に対するCTガイド下ラジオ波焼灼術(RFA)の治療成績
- (公社)日本医学放射線学会, 2021年, 日本医学放射線学会秋季臨床大会抄録集, 57th, S400 - S400, 日本語慢性血栓塞栓性肺高血圧症患者における,重回帰分析を用いた肺動脈平均圧推定についての検討
- (公社)日本医学放射線学会, 2020年10月, 日本医学放射線学会秋季臨床大会抄録集, 56回, S147 - S147, 日本語鈍的外傷による鎖骨下動脈損傷に対するステントグラフト内挿術の1例
- (公社)日本医学放射線学会, 2020年10月, 日本医学放射線学会秋季臨床大会抄録集, 56回, S147 - S147, 日本語鈍的外傷による鎖骨下動脈損傷に対するステントグラフト内挿術の1例
- (一社)日本インターベンショナルラジオロジー学会, 2020年08月, 日本インターベンショナルラジオロジー学会雑誌, 35(Suppl.) (Suppl.), 173 - 173, 日本語IVR初学者による中心静脈ポート留置術 内頸静脈vs鎖骨下静脈アプローチの比較検討
- (一社)日本インターベンショナルラジオロジー学会, 2020年08月, 日本インターベンショナルラジオロジー学会雑誌, 35(Suppl.) (Suppl.), 288 - 288, 日本語マイクロスフィアを用いた子宮筋腫塞栓術後のUFS-QOLによるQOL評価
- (一社)日本インターベンショナルラジオロジー学会, 2020年08月, 日本インターベンショナルラジオロジー学会雑誌, 35(Suppl.) (Suppl.), 243 - 243, 英語Spectral Detector CTによるEVAR後のendoleak検出能の検討 iodine no water imagesの有用性(Detectability of endoleaks after EVAR by spectral detector computed tomography)
- (公社)日本医学放射線学会, 2020年02月, 日本医学放射線学会学術集会抄録集, 79回, S156 - S156, 英語検出器スペクトラルCTによる血管内動脈瘤修復術後のエンドリークの検出可能性 iodine-no-water画像の有用性(Detectability of endoleaks after endovascular aneurysm repair on spectral detector computed tomography: utility of iodine-no-water images)
- (公社)日本医学放射線学会, 2020年02月, 日本医学放射線学会学術集会抄録集, 79回, S199 - S199, 英語Spectral detector CTによるS状結腸捻転による腸管虚血の評価(Evaluation of bowel ischemia caused by sigmoid volvulus using spectral detector computed tomography)
- (一社)日本インターベンショナルラジオロジー学会, 2019年05月, 日本インターベンショナルラジオロジー学会雑誌, 34(Suppl.) (Suppl.), 394 - 394, 日本語Paget-Schroetter syndromeに伴う上肢静脈血栓に対する血管内治療の一例
- (一社)日本インターベンショナルラジオロジー学会, 2019年05月, 日本インターベンショナルラジオロジー学会雑誌, 34(Suppl.) (Suppl.), 257 - 257, 日本語Chronic kidney disease(CKD)患者に対し炭酸ガス造影を用いたEVARの検討
- (公社)日本医学放射線学会, 2019年02月, Japanese Journal of Radiology, 37(Suppl.) (Suppl.), 34 - 34, 日本語白血病患児の治療経過中に発症した急性脳症の2例
- (一社)日本インターベンショナルラジオロジー学会, 2018年11月, IVR: Interventional Radiology, 33(3) (3), 319 - 319, 日本語巨大肺仮性動脈瘤に対してコイルおよびvascular plugで塞栓術を行った1例
- (公社)日本小児科学会, 2017年05月, 日本小児科学会雑誌, 121(5) (5), 901 - 901, 日本語硬膜下膿瘍を呈した18trisomyの1例
- (一社)日本小児血液・がん学会, 2015年10月, 日本小児血液・がん学会雑誌, 52(4) (4), 303 - 303, 日本語
- (公社)日本小児科学会, 2015年03月, 日本小児科学会雑誌, 119(3) (3), 628 - 628, 日本語急速な経過で悪化し、膿瘍形成した眼窩蜂窩織炎の1例
- (公社)日本小児科学会, 2015年02月, 日本小児科学会雑誌, 119(2) (2), 412 - 412, 日本語胸痛で発症した縦隔成熟奇形腫破裂の一例
- (一社)日本小児血液・がん学会, 2014年10月, 日本小児血液・がん学会雑誌, 51(4) (4), 298 - 298, 日本語
- (NPO)日本小児血液・がん学会・(NPO)日本小児がん看護学会・(公財)がんの子供を守る会, 2013年11月, 日本小児血液・がん学会学術集会・日本小児がん看護学会・公益財団法人がんの子どもを守る会公開シンポジウムプログラム総会号, 55回・11回・18回, 283 - 283, 日本語寛解導入療法後に結腸穿孔・巨大胃潰瘍を併発した難治性急性リンパ性白血病のダウン症候群の1例
- (一社)日本病理学会, 2007年02月, 日本病理学会会誌, 96(1) (1), 357 - 357, 日本語腺癌と絨毛上皮癌が共存した胆嚢癌の一剖検例
- ECR 2023 - European Congress of Radiology, 英語Construction of Prediction Models of Pericardial Fat Volume from Chest X-ray by Combined Three Different Deep Learning Methodsポスター発表
