Anirudh & Associates is an AI advisory and research lab.
Journal Careers HomeAnirudh & Associates is an independent advisory and research lab working on applied artificial intelligence. The practice exists because the distance between a model that demonstrates well and a system that holds up under real use is wider than most teams expect, and closing it is a specific discipline.
We take a small number of engagements at a time. That is a deliberate constraint: the problems worth taking are the ones where the answer is not known at the outset, and those cannot be run in parallel without becoming shallow.
Most teams arriving at an AI problem are not short of ideas. They are short of a way to tell which ideas are worth the next quarter. Advisory work begins there — with the decision, not the technology.
In practice that means mapping the problem to something measurable, establishing what a competent baseline actually scores, and being direct about where a language model is the wrong instrument. A good advisory engagement often ends with a smaller build than the one originally proposed.
Alongside client work the lab maintains an independent research strand. The questions are the ones that recur across engagements and never quite get answered inside any single one: how to measure quality on tasks with no clean ground truth, how retrieval and reasoning trade against each other under latency budgets, and where agentic systems fail in ways their traces do not reveal.
Results are published as they arrive, including the negative ones. Methods that did not work are the cheapest thing we can give away and often the most useful.
The method is unglamorous and it is the whole thing: define the task precisely, build the smallest honest evaluation of it, establish a baseline, then change one thing at a time.
Teams tend to skip the second step because it produces no visible progress. It is also the step that determines whether every subsequent decision is informed or guessed. An evaluation set that took two days to assemble will outlive three model migrations.
Evaluation is treated as a first-class artefact rather than a reporting obligation. That means version-controlled datasets, graders whose agreement with human judgement has itself been measured, and a clear account of what the numbers do not cover.
The failure we see most often is an evaluation suite that improves steadily while the product does not. Usually the suite has drifted toward what is easy to score. Keeping it honest is continuous work.
A model is one component in a system that also has caching, retries, fallbacks, rate limits, a cost ceiling, and a set of behaviours it must never exhibit. Most of what determines whether users trust the result happens in that surrounding machinery.
We work on the whole of it: the data path into the model, the evaluation path around it, and the operational path that tells you when something has quietly degraded in production.
Work generally takes one of three shapes. A diagnostic is a short, fixed-scope review of an existing system ending in a written assessment and a ranked set of recommendations. A build is hands-on delivery of a specific capability, run to a defined evaluation target. A standing advisory is recurring time with a team that can already build but wants a second opinion held to the same standard as the first.
Every engagement is scoped to a question that can be answered. If we cannot state what would count as an answer, that is a sign the engagement is not ready to start.
Measure before building. Prefer the smaller system that can be reasoned about. Report the result that was obtained rather than the one that was hoped for. Leave the client able to continue without us — an engagement that creates a dependency has failed at something, however good the deliverable.
And be willing to say that the problem does not need this. That answer is occasionally the most valuable one available, and it is the one an incentive-aligned vendor will never give.
The best starting point is a description of the problem in plain language, what you have tried, and what a good outcome would look like six months out. No brief required, and no expectation that the problem is already well formed — clarifying it is part of the work.
Write to apupneja2002@gmail.com. Everything is read.