From Volume to Value: Preference-Aligned Memory Construction for On-Device RAG
AuthorsChangmin Lee, Jaemin Kim, Taesik Gong
Resources
This paper shows how an on-device AI assistant can store less, remember better, and answer in ways that match your preferences using a compact preference-aligned memory system.
Key results
Across PrefWiki, PrefRQ, PrefELI5, and PrefEval, EPIC improves preference-following accuracy by this average margin over the best baseline, including NV-Embed-v2.
EPIC reduces the on-device indexing memory footprint by this factor by storing only preference-relevant items and compact instructions instead of raw traces.
EPIC shortens retrieval latency by this factor compared with the best-performing baseline while keeping retrieval as a single FAISS search over a smaller instruction index.
In the on-device streaming test on Jetson Orin Nano 8GB, EPIC keeps its memory footprint below 1 MB.
In the same on-device experiment, EPIC achieves this end-to-end retrieval latency per query, with steering overhead reported separately.
Preference-guided query steering adds only this amount of extra latency in the on-device retrieval pipeline.
What the paper found
This paper proposes EPIC, an on-device RAG memory construction framework that shifts personalization from storing all raw traces to storing only preference-relevant evidence. EPIC uses a three-stage pipeline: Semantic-Based Coarse Filtering with Contriever-style embeddings and cosine thresholding, Preference-Aligned Fine Verification with an LLM Decision Module and Instruction Generator, and Preference-Guided Query Steering that moves query embeddings toward the nearest preference vector before FAISS retrieval. The key novelty is instruction-centric indexing: each retained chunk is converted into a compact preference-conditioned instruction–item pair, so retrieval is guided by explicit usage directives rather than raw text. On four benchmarks—PrefWiki, PrefRQ, PrefELI5, and PrefEval—built from Wikipedia, Researchy Questions, Common Crawl, and LMSYS-Chat-1M, EPIC improves preference-following accuracy by 20.17 percentage points on average over the best baseline, including NV-Embed-v2, while reducing indexing memory by 2,404× and retrieval latency by 33.33×. In one on-device test on a Jetson Orin Nano 8GB, EPIC keeps memory under 1 MB and reaches 29.35 ms per query, with steering adding only 0.18 ms. Ablations show coarse filtering drives most of the memory reduction, fine verification is essential for accuracy, and query steering adds a further 0.78 to 4.03 percentage points without extra storage.
Original abstract
With the rapid emergence of personal AI agents based on Large Language Models (LLMs), implementing them on-device has become essential for privacy and responsiveness. To handle the inherently personal and context-dependent nature of real-world requests, such agents must ground their generation in device-resident personal context. However, under tight memory budgets, the core bottleneck is what to store so that retrieval remains aligned with the user. We propose EPIC (Efficient Preference-aligned Index Construction), which focuses on user preferences as a compact and stable form of personal context and integrates them throughout the RAG pipeline. EPIC selectively retains preference-relevant information from raw data and aligns retrieval toward preference-aligned contexts. Across four benchmarks covering conversations, debates, explanations, and recommendations, EPIC reduces indexing memory by 2,404 times, improves preference-following accuracy by 20.17 percentage points, and achieves 33.33 times lower retrieval latency over the best-performing baseline. In our on-device experiment, EPIC maintains a memory footprint under 1 MB with 29.35 ms/query latency in streaming updates.
Read the original paper