Lokesh L.K.S

Lokesh L.K.S

Software engineer · AI agents · backend systems

Open to software engineer and AI engineer roles · Bay Area or remote

I build AI agents and the systems that call AI models: the tools they use, the permissions they can't skip, and the services they run on.

M.S. Cybersecurity (Penn State), CCNA and Security+. I came to software through network security, which is why I treat model output as untrusted input. Away from the screen I skate and play chess.

build.logdocs/index.html
$ ./site_generator.exe=== Markdown Static Site Generator ===Copied: docs/style.cssProcessing page: index.mdPopulating blog database...Exported 19 blogs to docs/blogs.jsonGenerated: docs/index.html
the build that produced this page · SSGengine, my C++ static site generator
800 msJarvis Desk · decision budget before the LLM's first token 30 sGreenRoots · replay window on HMAC-signed service calls 5QUIC routing sim · RFC 9000 + QUIC-LB scenarios 700,774DDoS capstone · network flows modelled

Penn State M.S. Cybersecurity · Anna University (CEG) · Flower hackathon · Executable World hackathon

01 Now

I'm looking for software engineer and AI engineer roles: agents, backend, and developer tooling.

Experience

  1. AI/ML EngineerLokesh Agro Foods2025 – Jun 2026
  2. Infrastructure Security Engineer InterniMeritJun 2025 – Aug 2025
  3. Platform EngineerTvastrJul 2022 – Dec 2022
  4. AI Research InternAnna University (CEG)Feb 2023 – Jul 2024

Stack

Languages
Python, TypeScript, C++
Backend
FastAPI, PostgreSQL + pgvector, SQLite, server-sent events
AI
Claude API, Gemini, OpenAI Agents SDK, PyTorch
Tooling
Docker, CMake, GitHub Actions, Jupyter

.ts · jarvis-desk

Jarvis Desk

Executable World hackathon · Track 1 (AI Assistants)

A voice-first assistant with a two-speed brain. Every turn starts with a typed decide() call under an 800 ms budget that classifies intent and danger before the LLM produces a token. Permission tiers are enforced in code, not in the prompt: DANGER tools always stop for a human.

A web HUD and a Windows overlay speak the same streaming /chat protocol.

Details ↓

Flower Hub · @lokilks/second-brain

Second Brain

Turns one goal into milestones and scored decision cards from four specialist agents — planner, events, health, content. Nothing executes without an approved card, and every approval and skip feeds the next ranking; the loop closes through WhatsApp and Google Calendar.

Details ↓

02 Work

AI agents

Backend & systems

ML & security engineering

LLM evaluation

Lab 3 projects

Earlier research on model internals: auditing fine-tunes, interpretability and probes. Kept here for range.

  • .py Detecting secret loyalties in fine-tuned models Apart Research sprint · track 2 · Qwen2.5-7B-Instruct 8,700 generations

    Audited three fine-tunes of Qwen2.5-7B-Instruct for covert, weight-encoded loyalties across ~8,700 generations, 127 actors and six trigger classes. Model C was confirmed as the null four independent ways: byte-identical safetensors checksums, bitwise-identical activations at all 29 layers, 1000/1000 verbatim-identical generations at matched seeds, and zero entity deltas on stance probes.

    Models A and B were genuinely modified — activation divergence from base peaks at layers 10–12 — but no secret loyalty was detectable in either, under per-actor separability testing, next-token affinity probing, and causal activation patching that came back actor-blind. Reported as a rigorous negative result rather than stretched into a finding.

    The part worth pointing at: prefill elicitation extracted specific “confessions” from the known-clean base model at 22.9% — indistinguishable from A at 25.0% and B at 28.1%. A method that produces confessions from a model you can prove is clean cannot be trusted on one you can't.

    • 127 actors · 6 trigger classes
    • activation patching
    • white-box + black-box
    • A100-80GB
    • negative result, reported as one
  • .ipynb The modern interpretability stack, built from scratch SAE → transcoder → cross-layer transcoder → attribution graph GPT-2 small

    The full modern mech-interp ladder on GPT-2 Small, each rung implemented and measured: sparse autoencoders for which concepts are present, transcoders for what an MLP computes, cross-layer transcoders for computation spanning depth, and attribution graphs for the causal chain behind a single prompt.

    It ends where interpretability work should end — not at a plausible diagram, but at a test. The graph names the features it claims carry the indirect-object circuit; ablate exactly those and the model's answer collapses as predicted.

    • sparse autoencoders
    • transcoders
    • cross-layer transcoders
    • attribution graphs
    • IOI circuit
    • causal validation
  • .ipynb Linear and deception probes reading internal states the output denies d_model 3,584

    Activation extraction with manual PyTorch forward hooks, then linear probes fitted on the residual stream — 3,584 dimensions per token on 7B models — using contrast pairs and difference-of-means. When a linear classifier succeeds, the concept is approximately a direction in activation space, which hands you a vector you can manipulate rather than just a score.

    The interesting case is the deception probe: training a reader for internal states the model's own text output actively denies. Where in the stack a probe starts working is itself the result — one that succeeds at layer 0 is reading surface form, not meaning.

    • PyTorch forward hooks
    • contrast pairs
    • difference-of-means
    • deception probes

03 Writing

  1. When AI Agrees With You Just to Be Nice: Measuring Sycophancy in Llama-3.2-3B

    A few months ago, I asked Llama-3.2-3B whether the Earth is flat.

  2. When AI Shows Its Work—But Lies About How It Got There: Measuring CoT Faithfulness in Llama-3-8B

    I gave Llama-3-8B a logic puzzle and a nudge.

  3. Building a Temporal CNN for DDoS Detection

    The capstone started as a straightforward ML project: train a model to detect DDoS attacks.

All writing →