International Journal of Innovative Research in Engineering and Management
Year: 2026, Volume: 13, Issue: 3
First page : ( 125) Last page : ( 137)
Online ISSN : 2350-0557
DOI: 10.55524/ijirem.2026.13.3.15 |
DOI URL: https://doi.org/10.55524/ijirem.2026.13.3.15
This is an Open Access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0) (http://creativecommons.org/licenses/by/4.0)
Article Tools: Print the Abstract | Indexing metadata | How to cite item | Email this article | Post a Comment
Piyush Thapliyal , Purva Mundada, Mukthikka V, Gurpreet Singh
Current systems that generate content across multiple modes, like video and audio, often face issues such as missing objects, weak connections between meanings, and poor coordination between visual and sound elements. These problems often arise because visual and audio parts are created separately, leading to mismatches between what is shown and what is heard. To tackle these issues, this paper introduces SG-WORLD V2, a framework that focuses on planning first for semantic reasoning and coordination in structured world simulations. Instead of directly creating video or audio, SG-WORLD V2 builds a semantic blueprint first. It translates natural language into machine-readable forms using reasoning based on the Universal Scene Graph (USG), adds a Quantification Rule to keep track of requested items; and uses Deterministic Pre-Temporal Synchronization (DPTS) to align audio and visual events during planning. The system also includes a self-correcting mechanism to check for semantic consistency before generating outputs. The framework was tested with 54 different prompts that cover object arrangements, spatial connections, actions, and synchronized audio-video situations. The results showed strong performance in keeping track of objects, maintaining relationships, and creating logically aligned plans. Overall, SG-WORLD V2 offers a clear and dependable semantic planning method that can support future systems for generating video and audio while improving consistency and alignment between different modes.
MS Scholar, Endicott College of International Studies, Woosong University, Daejeon, Korea
No. of Downloads: 10 | No. of Views: 137
Nidhi Singh.
June 2026 - Vol 13, Issue 3
T. Srajan Kumar, Kandula Siri Chandana, Gavara Archana, Gundeti Dhanush Reddy, Jonna Madhu Reddy.
April 2026 - Vol 13, Issue 2
Devi Priya Gottumukkala, K. Mounika, S. J. Harivallika, J. Vinay Kumar, K. Surya Siddhu.
April 2026 - Vol 13, Issue 2
