TY - GEN
T1 - From Pixels to Portions
T2 - 1st International Conference on Data Science and Geoinformatics, ICDSG 2025
AU - Rizaluddin, Baghas
AU - Sari, Yuita Arum
AU - Adinugroho, Sigit
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Food leftovers in hospitals are a critical issue, directly impacting the quality of patient nutrition abd the efficiency of the hospital operating budget. This study proposes an automated framework to estimate the percentage of food leftovers from patient meals using deep learning-based image analysis. The proposed method utilizes the combination of Mask R-CNN model for object pixel-wise segmentation and Depth Anything V2 for monocular depth estimation, with the calculation of leftovers based on before and after object volumes. The dataset used in this study, referred to as Le-Food dataset, consists of manually captured food images in a controlled environment to maintain uniform conditions. Experimental results show that deeper backbones improve accuracy but require a higher amount of computational process, ResNet-50 + Depth Anything V2 VIT-L combination achieved the lowest Mean Absolute Error (MAE) of 10.516% with an inference time of 6.275s per pair. However, a lighter model, ResNet-50 + Depth Anything V2 VIT-S, offered a more practical compromise, producing a competitive MAE of 10.987% at 4.878s, the fastest combinations among all. These findings demonstrate that an objective, automated approach to monitoring food waste is viable. The proposed framework contributes to the improvement of hospital nutritional services and supports a data-driven solution in dietary planning for hospital patients. In addition, this approach can be generalized to other healthcare or food services where volume estimation is essential.
AB - Food leftovers in hospitals are a critical issue, directly impacting the quality of patient nutrition abd the efficiency of the hospital operating budget. This study proposes an automated framework to estimate the percentage of food leftovers from patient meals using deep learning-based image analysis. The proposed method utilizes the combination of Mask R-CNN model for object pixel-wise segmentation and Depth Anything V2 for monocular depth estimation, with the calculation of leftovers based on before and after object volumes. The dataset used in this study, referred to as Le-Food dataset, consists of manually captured food images in a controlled environment to maintain uniform conditions. Experimental results show that deeper backbones improve accuracy but require a higher amount of computational process, ResNet-50 + Depth Anything V2 VIT-L combination achieved the lowest Mean Absolute Error (MAE) of 10.516% with an inference time of 6.275s per pair. However, a lighter model, ResNet-50 + Depth Anything V2 VIT-S, offered a more practical compromise, producing a competitive MAE of 10.987% at 4.878s, the fastest combinations among all. These findings demonstrate that an objective, automated approach to monitoring food waste is viable. The proposed framework contributes to the improvement of hospital nutritional services and supports a data-driven solution in dietary planning for hospital patients. In addition, this approach can be generalized to other healthcare or food services where volume estimation is essential.
KW - computer vision
KW - depth estimation
KW - healthcare
KW - object segmentation
KW - volume calculation
UR - https://www.scopus.com/pages/publications/105038037154
U2 - 10.1109/ICDSG67714.2025.11381375
DO - 10.1109/ICDSG67714.2025.11381375
M3 - Conference contribution
AN - SCOPUS:105038037154
T3 - 2025 1st International Conference on Data Science and Geoinformatics, ICDSG 2025
SP - 79
EP - 84
BT - 2025 1st International Conference on Data Science and Geoinformatics, ICDSG 2025
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 26 November 2025 through 28 November 2025
ER -