
Six months ago, our ai code assistant had a problem。
Ask it:
"what services would be affected if i modified the payment stream?"
It returns confidently to a partially correct answer。
It's not about llm。
The problem is the context。
Our warehouse contains over half a million lines of code, distributed among api, micro-services, event consumers, time assignments, shared libraries and infrastructure codes。
Traditional vector searches can find relevant documents。
However, it was difficult to understand how all the content was connected。
So we turned the code library from document collections to knowledge mapping。
The result is an ai system that answers architecture questions, performs impact analysis, identifies hidden dependencies and helps engineers to navigate large code libraries more accurately。
Here's how we build it。
One, why ai struggles in a large code library
Most ai coding tools rely heavily on semantic searches。
The process was broadly as follows:
User problems
zenium
vector search
zenium
relevant documents
zenium
llm
zenium
answer
This has been of extraordinary benefit to small projects。
However, large enterprise systems are different。
Imagine the following chain of dependence:
United services
zenium
billing service
zenium
paymentgateway
zenium
stripeadapter
Now ask:
What would stripeadapter destroy
Vector search may search:
Stripeadapter. Ts
paymentgateway. Ts
But it could be completely missed:
Billing service
united services
orderprocessing
refundworkflow
The problem is the similarity of vector search understanding。
It does not understand relationships naturally。
The software system is built on relationships。
Function。
Category of succession。
Service release event。
Consumer subscription event。
Module。
Without these relationships, ai sees code fragments rather than architecture。
2. Consider codes as maps
Once we step back, the solution becomes obvious。
The library is naturally a map。
Funaction
zenium
calls
funaction
class. Zenium
inherits
class. Service
zenium
i'm sorry. Event. CoI'm sorry. Zenium
i'm sorry. Event
Each node represents meaningful content。
Example:
File
class. Method
funaction
service
event. DatabI'm sorry. Api endpoint
queue
Each side represents a relationship。
Example:
Calls
i..Mports
inherits
publishes
subscribes
reads
writes
depends on
Once you model the warehouse in this way, it becomes much easier to answer structural questions。
3. Extract structures from warehouses
The first step is to build a solver。
For typesCript service, we use ts-morph to run through the abstract syntax tree (ast)。
IMport {project} from “ts-morph”;
coNst project = new project();
(b) project. Addsourcefilesatpaths (“src*. Ts”);
for (co)Nst file of project. Getsourcefiles() {
coI'm sorry= file. GetiMportdeclarations();
iOther organiser
i don't know. I'm sorryPleasename(),
this post is part of our special coverage global voices 2011.
"i."Mports"
i'm not sure. I'm not sure.
♪ i'm sorry ♪
This gives us a relationship:
United services
i..Mports
billing service
i..Mports
paymentgateway
Next, we extract:
Target is not index code。
The target is mapping relationships。
Building knowledge maps spectrum
After extraction, we store them in neo4j。
Example:
Create
(a: service {name: "use"))
-[:depends on]-
(b: service {name: "biling service"})
Another example:
Create
(a: service {name: "biling service"})
-[:calls] ->
(b: service {name: "paymentgateway"})
The figure began to grow very quickly。
We no longer have thousands of isolated documents but have:
20,000+node
85,000+ relationships
That's where things get interesting。
5. Impact analysis becomes simple
Previously, it was painful to answer that question:
What would happen if i changed payment gateway
Engineers manually check:
Sometimes it takes hours。
Here's the picture:
Match p=(n)-[*]->(m)
what's your name? Return p
The answer immediately emerged。
The figure shows each downstream dependency relationship。
This is one of the most valuable capabilities we have built。
6. Combining the chart search with llm
The figure itself is useful。
The real breakthrough happens when we connect it to llm。
Structure:
User problems
zenium
figure query
zenium
related subgraphs
zenium
code search
zenium
llm
zenium
answer
Assuming the developers ask:
Which services are affected by the re-testing of payments
The workflow becomes:
Step 1: identification of entities。
♪ payment returns ♪
paymentservice
retryprocessor
Step 2: query chart relationships。
Paymentservice
zenium
retryprocessor
zenium
notificatioNo, no, no, no. Zenium
billing service
Step 3: search for actual codes。
I don't know. I'm sorry. It's not gonna happen, billing
Step 4: send only relevant context to llm。
Instead of storing hundreds of files into context windows, we provide a highly connected subset of the warehouse。
The quality of answers has improved significantly。
Seven
We built it to help ai。
Surprisingly, humans began to use it more frequently than ai。
Engineers started asking questions like:
Previously these answers needed tribal knowledge。
You can check them now。
It became a living chart。
8 beyond rag
Most of the teams that tried to improve the ai coding assistant focused on better retrieval。




