TY - GEN
T1 - From Visual Perception to Context-Aware Instructions
T2 - 2025 International Conference on AI-Driven Business Transformation and Data Science Innovation, ICBTDS 2025
AU - Tama, Nazura Wirayuda
AU - Yudistira, Novanto
AU - Manurung, Daniel Geoffrey
AU - Sulaiman, Sarina
AU - Yong, Pang Yee
AU - Muchtar, Farkhana
AU - Othman, Nur Zuraifah Syazrah
AU - Ismail, Nor Azman
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2026/1/22
Y1 - 2026/1/22
N2 - Modern navigation systems often stop at visual perception, providing raw detections without translating them into actionable guidance for drivers. We present a practical, deployable navigation assistance framework that tightly couples real-time object detection with Large Language Models (LLMs) to produce fluent, context-aware driving instructions. Unlike conventional Advanced Driver Assistance Systems (ADAS), our approach introduces a structured intermediate representation that encodes detector outputs, such as traffic sign identities and locations, into a compact, machine-readable format before prompting the LLM. This design improves controllability, reduces irrelevant generation, and enables adaptation to dynamic roadway and traffic contexts. Evaluated on the 21-class Traffic Sign in Indonesia Dataset, our system achieves state-of-the-art detection performance and introduces the Feasibility Score, a multi-criteria human evaluation metric that captures relevance, coherence, completeness, fluency, and specificity of generated instructions. Experiments across multiple LLM configurations demonstrate that coupling structured perception with LLM-based reasoning produces guidance rated as clearer, more specific, and more context-relevant than perception-only or naive LLM baselines. These results position our framework as a concrete step toward next-generation, human-centered navigation assistance that bridges the gap between visual recognition and actionable driver communication.
AB - Modern navigation systems often stop at visual perception, providing raw detections without translating them into actionable guidance for drivers. We present a practical, deployable navigation assistance framework that tightly couples real-time object detection with Large Language Models (LLMs) to produce fluent, context-aware driving instructions. Unlike conventional Advanced Driver Assistance Systems (ADAS), our approach introduces a structured intermediate representation that encodes detector outputs, such as traffic sign identities and locations, into a compact, machine-readable format before prompting the LLM. This design improves controllability, reduces irrelevant generation, and enables adaptation to dynamic roadway and traffic contexts. Evaluated on the 21-class Traffic Sign in Indonesia Dataset, our system achieves state-of-the-art detection performance and introduces the Feasibility Score, a multi-criteria human evaluation metric that captures relevance, coherence, completeness, fluency, and specificity of generated instructions. Experiments across multiple LLM configurations demonstrate that coupling structured perception with LLM-based reasoning produces guidance rated as clearer, more specific, and more context-relevant than perception-only or naive LLM baselines. These results position our framework as a concrete step toward next-generation, human-centered navigation assistance that bridges the gap between visual recognition and actionable driver communication.
KW - LLM
KW - Navigation Assistance
KW - Object Detection
UR - https://www.scopus.com/pages/publications/105030319232
U2 - 10.1145/3786554.3786576
DO - 10.1145/3786554.3786576
M3 - Conference contribution
AN - SCOPUS:105030319232
T3 - Proceedings of 2025 International conference on AI-Driven Business Transformation and Data Science Innovation, ICBTDS 2025
SP - 138
EP - 144
BT - Proceedings of 2025 International conference on AI-Driven Business Transformation and Data Science Innovation, ICBTDS 2025
PB - Association for Computing Machinery, Inc
Y2 - 14 November 2025 through 16 November 2025
ER -