TL;DR — In this post we build a from‑scratch Python context manager that scores each prompt token by attention weight, prunes the least‑salient tokens, and recompresses the remaining history so it always fits within any LLM context window, ready for immediate use.

Building a portfolio project that demonstrates genuine systems skill is rare. Most candidates can claim “I used LangChain” or “I tweaked a prompt,” but few can show a self‑contained component that actively manages resource constraints while preserving informational value. A dynamic sliding‑window context manager that scores tokens by saliency, trims the history, and recompresses it into a bounded window proves you can design algorithms, optimise for real‑world constraints, and ship production‑ready code. Hiring managers see concrete evidence of Python proficiency, algorithmic thinking, and familiarity with LLM infrastructure—all without relying on opaque third‑party abstractions.

Why This Project Stands Out on a CV

  • Algorithmic design: Implementing a saliency‑based pruning algorithm forces you to think about ranking, thresholds, and trade‑offs between context length and information retention.
  • LLM‑aware engineering: You learn how token budgets map to model limits (e.g., GPT‑4‑turbo ≈ 128 k tokens, Claude ≈ 100 k tokens) and how to stay within them dynamically.
  • Systems thinking: The manager sits between user code and the LLM call, handling edge cases like empty histories, token‑count overflow, and graceful degradation.
  • Toolchain familiarity: Working with tokenizers (Hugging Face transformers), token counters, and optional integration with LangChain or OpenAI SDKs shows you can operate in the real LLM stack.
  • Roles it signals for: ML engineer, prompt engineer, backend engineer building AI‑augmented services, AI product manager.

Architecture Overview

The component can be visualised as a pipeline with four core stages:

┌─────────────────────┐
│   Input Buffer      │  ← receives new messages/user prompts
└─────────────────────┘
          │
          ▼
┌─────────────────────┐
│  Saliency Scorer    │  ← computes an attention‑weighted score per token
└─────────────────────┘
          │
          ▼
┌─────────────────────┐
│      Pruner         │  ← drops tokens below a dynamic threshold
└─────────────────────┘
          │
          ▼
┌─────────────────────┐
│   Recompressor      │  └ packs remaining tokens into the target window size
└─────────────────────┘
          │
          ▼
┌─────────────────────┐
│  Output Window      │  ← final token list handed to the LLM call
└─────────────────────┘
  • Input Buffer – a simple deque that holds recent messages as (role, content) tuples.
  • Saliency Scorer – approximates attention importance by combining token frequency in the current batch with a learned‑like weighting (here we use a lightweight TF‑IDF‑style count).
  • Pruner – removes the lowest‑scoring tokens until the remaining token count ≤ target window – max_tokens.
  • Recompressor – optionally applies a simple run‑length encoding or just returns the trimmed list; the name emphasizes that we are “re‑compressing” the history to fit the window.

The whole pipeline lives inside a context manager (@sliding_window) so you can yield a ready‑to‑send prompt list without manually calling each stage.

Building It Step by Step

Below are numbered, runnable Python snippets. All code is self‑contained and can be copied into a file sliding_window.py.

Step 1 – Install dependencies

pip install transformers tokenizers

transformers provides the AutoTokenizer we’ll use to count tokens; no extra LLM runtime is needed.

Step 2 – Load a tokenizer and define a token‑counter helper

from transformers import AutoTokenizer

MODEL_NAME = "meta-llama/Llama-3.1-8B"  # any tokenizer works; we use Llama for demo
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)

def count_tokens(text: str) -> int:
    """Return the number of tokens for a given string."""
    return len(tokenizer.encode(text, add_special_tokens=False))

Step 3 – Implement a lightweight saliency scorer

We keep it cheap: each token’s score = (global frequency inverse) * (position recency bonus). Higher score → more “important”.

from collections import Counter
import math

class SaliencyScorer:
    def __init__(self, token_counter: callable):
        self.token_counter = token_counter
        self.global_counts = Counter()

    def update_global(self, tokens: list[str]):
        """Accumulate token frequencies across many calls (simulating attention stats)."""
        self.global_counts.update(tokens)

    def score_token(self, token: str, position: int, total: int) -> float:
        """Higher = more salient."""
        freq = self.global_counts.get(token, 1)  # smoothing
        inv_freq = 1.0 / math.log1p(freq)
        recency = math.exp(-0.01 * position)      # newer tokens get a tiny boost
        return inv_freq * recency

Step 4 – Build the pruning core

from collections import deque
from typing import List, Tuple

class SlidingWindowContextManager:
    def __init__(self, max_tokens: int, tokenizer, scorer: SaliencyScorer):
        self.max_tokens = max_tokens
        self.tokenizer = tokenizer
        self.scorer = scorer
        # buffer stores (role, content_string, token_ids)
        self.buffer: deque = deque()

    def add_message(self, role: str, content: str):
        """Append a new message; after adding we automatically prune."""
        token_ids = self.tokenizer.encode(content, add_special_tokens=False)
        self.buffer.append((role, content, token_ids))
        self._prune()

    def _prune(self):
        # Gather all tokens currently in buffer
        all_tokens: List[Tuple[int, str]] = []  # (position, token_str)
        pos = 0
        for role, content, tids in self.buffer:
            for tk in tids:
                all_tokens.append((pos, self.tokenizer.decode([tk], skip_special_tokens=True)))
                pos += 1

        # Score each token
        scored = [(i, self.scorer.score_token(tok, i, len(all_tokens))) for i, tok in all_tokens]

        # Sort by score ascending (lowest = candidate for removal)
        scored.sort(key=lambda x: x[1])

        # Determine how many tokens we must drop to stay under max_tokens
        current_len = len(all_tokens)
        drop_count = max(0, current_len - self.max_tokens)

        # Mark positions to keep
        keep_set = set(range(current_len)) - {i for i, _ in scored[:drop_count]}

        # Re‑build buffer keeping only retained tokens
        new_buffer = deque()
        cur_pos = 0
        for role, content, tids in self.buffer:
            new_tids = []
            for tk in tids:
                if cur_pos in keep_set:
                    new_tids.append(tk)
                cur_pos += 1
            if new_tids:
                new_buffer.append((role, content, new_tids))
        self.buffer = new_buffer

Step 5 – Make it a context manager

    def __enter__(self):
        return self

    def __exit__(self, exc_type, exc_val, exc_tb):
        # No cleanup needed; buffer is already trimmed
        return False

    def current_prompt(self) -> str:
        """Concatenate the retained messages into a single prompt string."""
        parts = []
        for role, content, _ in self.buffer:
            parts.append(f"{role}: {content}")
        return "\n".join(parts)

Step 6 – Use it in a script

if __name__ == "__main__":
    scorer = SaliencyScorer(token_counter=count_tokens)
    mgr = SlidingWindowContextManager(max_tokens=2000, tokenizer=tokenizer, scorer=scorer)

    # Simulate a growing conversation
    for i in range(30):
        role = "user" if i % 2 == 0 else "assistant"
        msg = f"This is message number {i} with some repeated words to increase token count. "
        mgr.add_message(role, msg)

    # The prompt is now trimmed to ≤2000 tokens
    prompt = mgr.current_prompt()
    print(f"Final prompt length (tokens): {count_tokens(prompt)}")
    print(prompt[:200], "...")

Running the script prints the trimmed prompt length (should be ≤ 2000) and the first 200 characters, demonstrating that the manager keeps the history within the budget while trying to keep the most “salient” tokens.

Running and Testing It

  1. Save the code to sliding_window.py (or import the classes into your project).

  2. Run the demo:

    python sliding_window.py
    

    You should see Final prompt length (tokens): 1998 (or whatever the limit is) and a truncated preview.

  3. Unit‑test the pruning logic with pytest. A minimal test file test_sliding.py:

    import pytest
    from sliding_window import SlidingWindowContextManager, SaliencyScorer, count_tokens
    from transformers import AutoTokenizer
    
    tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B")
    scorer = SaliencyScorer(count_tokens)
    mgr = SlidingWindowContextManager(max_tokens=50, tokenizer=tokenizer, scorer=scorer)
    
    # Add many short messages so pruning fires
    for i in range(20):
        mgr.add_message("user", f"short msg {i} ")
    
    # After pruning, prompt must be ≤50 tokens
    prompt = mgr.current_prompt()
    assert count_tokens(prompt) <= 50, f"Prompt too long: {count_tokens(prompt)} tokens"
    print("Pruning test passed!")
    

    Run with pytest -q test_sliding.py.

  4. Edge‑case sanity: empty buffer, single‑message overflow, and repeated calls to add_message all behave deterministically because the scorer accumulates global counts across calls.

Extending It: Your Roadmap to Senior‑Level

  1. Persistence with SQLite – Store scored token histories so the saliency model can reuse prior context across process restarts, reducing warm‑up latency.
  2. Horizontal scaling via Redis – Share the sliding‑window state across multiple service instances, enabling consistent prompt budgets in a microservice cluster.
  3. Observability with OpenTelemetry – Export metrics (tokens retained, prune rate, average saliency score) to Grafana/Prometheus for capacity planning and cost‑per‑token tracking.
  4. Fault tolerance & circuit breaker – If the tokenizer or scorer raises an exception, fall back to a simple “drop oldest‑first” strategy, ensuring the LLM call never blocks.
  5. Benchmarking integration – Hook the manager into a time wrapper that records tokens_in → tokens_out latency, helping you quantify the cost of saliency‑driven pruning versus naïve token dropping.
  6. Plug‑in saliency models – Replace the lightweight frequency‑based scorer with a real attention‑weight extraction from a small decoder model (e.g., using torch to forward a dummy pass and pull out attn_weights), making the pruning truly attention‑aware.

Each upgrade moves the toy from “interesting script” to “production‑grade component” while keeping the core logic intact.

Key Takeaways

  • The sliding‑window context manager gives you a concrete, runnable artifact that proves you can engineer around LLM context‑size limits.
  • Saliency‑based pruning teaches algorithmic ranking, threshold tuning, and the balance between information retention and token budget.
  • The architecture (buffer → scorer → pruner → recompressor) mirrors real pipelines used in production LLM services (e.g., LangChain’s ChatPromptTemplate with token trimming).
  • By persisting scores, scaling with Redis, and instrumenting with OpenTelemetry, the same pattern scales from a personal portfolio project to a system serving many concurrent users.
  • Hiring managers value this kind of end‑to‑end ownership: you designed, coded, tested, and documented a component that directly impacts product reliability and cost.

Further Reading