Latency-Aware Decision Intelligence: Neural Architecture Search for Auto-Optimizing ML Inference Pipelines Under SLA Constraints in Retail Analytics

Authors

  • Krishna kanth Thottempudi Infosys Limited, USA
  • Radhika Kande Premier Inc, USA
  • Chaithanya Kotla Devops and Cloud lead, State of Maryland, USA

Keywords:

neural architecture search, latency-aware inference, SLA optimization, retail analytics, inference pipeline optimization, decision intelligence.

Abstract

Retail analytics increasingly depends on ML inference pipelines that must deliver predictions within strict service-level time limits while preserving useful decision quality. Neural architecture search and deployment-aware optimization have improved serving efficiency, yet many retail systems still optimize for offline accuracy first and address latency only later through compression or tuning. The main gap is the lack of frameworks that treat latency, throughput, and SLA compliance as direct search objectives during pipeline design. This matters because SLA violations in recommendation, forecasting, and real-time retail decision services can reduce responsiveness and weaken operational reliability under peak load. This article presents a latency-aware decision intelligence framework based on neural architecture search for auto-optimizing ML inference pipelines under strict SLA constraints in retail analytics. The results show reduced SLA violation rates across search iterations and stronger inference accuracy across latency levels than baseline and manually optimized pipelines. The study shows that SLA-aware search can produce retail inference pipelines that are both deployable and analytically effective.

Downloads

Published

2024-12-17

Issue

Section

Articles