Skip to main navigation Skip to search Skip to main content

From Visual Perception to Context-Aware Instructions: Integrating Object Detection and LLMs for Navigation Assistance

  • Nazura Wirayuda Tama
  • , Novanto Yudistira*
  • , Daniel Geoffrey Manurung
  • , Sarina Sulaiman
  • , Pang Yee Yong
  • , Farkhana Muchtar
  • , Nur Zuraifah Syazrah Othman
  • , Nor Azman Ismail
  • *Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Modern navigation systems often stop at visual perception, providing raw detections without translating them into actionable guidance for drivers. We present a practical, deployable navigation assistance framework that tightly couples real-time object detection with Large Language Models (LLMs) to produce fluent, context-aware driving instructions. Unlike conventional Advanced Driver Assistance Systems (ADAS), our approach introduces a structured intermediate representation that encodes detector outputs, such as traffic sign identities and locations, into a compact, machine-readable format before prompting the LLM. This design improves controllability, reduces irrelevant generation, and enables adaptation to dynamic roadway and traffic contexts. Evaluated on the 21-class Traffic Sign in Indonesia Dataset, our system achieves state-of-the-art detection performance and introduces the Feasibility Score, a multi-criteria human evaluation metric that captures relevance, coherence, completeness, fluency, and specificity of generated instructions. Experiments across multiple LLM configurations demonstrate that coupling structured perception with LLM-based reasoning produces guidance rated as clearer, more specific, and more context-relevant than perception-only or naive LLM baselines. These results position our framework as a concrete step toward next-generation, human-centered navigation assistance that bridges the gap between visual recognition and actionable driver communication.

Original languageEnglish
Title of host publicationProceedings of 2025 International conference on AI-Driven Business Transformation and Data Science Innovation, ICBTDS 2025
PublisherAssociation for Computing Machinery, Inc
Pages138-144
Number of pages7
ISBN (Electronic)9798400722233
DOIs
Publication statusPublished - 22 Jan 2026
Event2025 International Conference on AI-Driven Business Transformation and Data Science Innovation, ICBTDS 2025 - Bandung, Indonesia
Duration: 14 Nov 202516 Nov 2025

Publication series

NameProceedings of 2025 International conference on AI-Driven Business Transformation and Data Science Innovation, ICBTDS 2025

Conference

Conference2025 International Conference on AI-Driven Business Transformation and Data Science Innovation, ICBTDS 2025
Country/TerritoryIndonesia
CityBandung
Period14/11/2516/11/25

Keywords

  • LLM
  • Navigation Assistance
  • Object Detection

Fingerprint

Dive into the research topics of 'From Visual Perception to Context-Aware Instructions: Integrating Object Detection and LLMs for Navigation Assistance'. Together they form a unique fingerprint.

Cite this