Adding AI wasn’t the difficult part. Deciding where AI belongs in the application architecture was.
Adding an LLM API to an application is relatively straightforward.
The harder questions begin after the first API call works.
Should AI run inside the user’s HTTP request? What happens when the provider is slow? Which AI operations should be asynchronous? What belongs in React state versus server state? Should an AI-generated support reply be sent automatically? What happens when a background job retries after partially completing?
I ran into these questions while building an AI-powered customer support application with React, TypeScript, Node.js, PostgreSQL, Redis, BullMQ and OpenAI.
The interesting part of the project wasn’t the prompt.
It was deciding where AI should—and should not—sit inside the system.
This article walks through that architecture, shows selected code from the actual implementation, and discusses several things I would change before scaling the design further.
What the Application Does
The application revolves around support tickets and conversations.
It has three roles:
- Customers create tickets, read their own tickets and add messages.
- Agents work with assigned tickets, update their status and request AI-generated reply suggestions.
- Admins have broader ticket access and assignment capabilities.
AI assists the workflow in two different ways.
First, when a new message is added, background processing analyzes the ticket conversation to generate:
- a ticket summary,
- sentiment,
- priority.
Second, an agent or admin can explicitly request an AI-generated reply suggestion.
Those two AI workloads look similar at first—they both call an LLM—but they have very different execution requirements.
That distinction shaped much of the architecture.
High-Level Architecture
At a high level, the system looks like this:
┌─────────────────────────┐
│ React Client │
│ │
│ Auth / Tickets / │
│ Messages / AI Replies │
└────────────┬────────────┘
│
REST API
│
┌────────────▼────────────┐
│ Node.js / Express │
│ │
│ Auth → Controller → │
│ Service → Repository │
└───────┬─────────┬───────┘
│ │
│ │ enqueue
│ ▼
┌───────▼───┐ ┌──────────────┐
│PostgreSQL │ │ Redis/BullMQ │
└───────────┘ └──────┬───────┘
│
▼
┌─────────────┐
│ AI Worker │
│ │
│ Summary │
│ Sentiment │
│ Priority │
└──────┬──────┘
│
LLM provider
The application is deliberately not split into many network services.
The API is a modular monolith, while AI background execution runs in a separate worker process.
That gives me an important boundary without introducing a distributed-service architecture prematurely:
HTTP request handling and background AI processing have independent execution lifecycles.
Frontend State: Don’t Put Everything in One Store
One frontend decision was to separate state based on who owns it.
I use three broad categories:
- authentication/session state → React Context,
- remote server state → TanStack Query,
- transient UI state → local component state.
For example, tickets are server-owned resources, so their fetching and cache lifecycle belong to TanStack Query:
export function useTickets(params?: TicketListParams) {
return useQuery({
queryKey: ticketKeys.list(params),
queryFn: () => getTickets(params),
});
}
export function useTicket(ticketId: string) {
return useQuery({
queryKey: ticketKeys.detail(ticketId),
queryFn: () => getTicketById(ticketId),
enabled: Boolean(ticketId),
});
}
export function useCreateTicket() {
const queryClient = useQueryClient();
return useMutation({
mutationFn: (payload: CreateTicketRequest) => createTicket(payload),
onSuccess: () => {
queryClient.invalidateQueries({ queryKey: ticketKeys.lists() });
},
});
}
The important part here isn’t the library.
It’s the ownership model.
A ticket fetched from the API is not fundamentally client state. The server remains authoritative.
TanStack Query handles fetching, caching and invalidation, while component-local state handles things such as editable drafts and temporary UI interactions.
This prevents a common React architecture problem: turning one global store into a mirror of the backend.
The Backend Is a Modular Monolith
The backend follows a layered flow that is roughly:
Route
↓
Authentication / Authorization / Validation
↓
Controller
↓
Service
↓
Repository
↓
Prisma
↓
PostgreSQL
This gives features such as tickets and messages clear boundaries without requiring each feature to become an independent service.
The service layer owns application workflows.
Repositories isolate most persistence operations.
Controllers remain focused primarily on translating HTTP requests and responses.
AI is a partial exception because some worker-side operations interact more directly with Prisma, but the overall HTTP architecture follows this separation.
For the current scope, I prefer this to creating multiple microservices.
A service boundary should solve an actual deployment, ownership or scaling problem—not exist simply because the system has multiple features.
Why I Moved AI Analysis Out of the Request Path
Consider what happens when a customer adds a message.
The important user-facing operation is:
Persist my message successfully.
Summary, sentiment and priority enrichment are useful, but the user should not have to wait for several LLM operations to complete before the message request can finish.
The service therefore persists the message and publishes a BullMQ job:
export const createMessageService = async (data: CreateMessageInput) => {
await getAuthorizedTicketService(data.ticketId, data.accessContext);
const message = await createMessageRepository({
content: data.content,
ticketId: data.ticketId,
senderId: data.senderId,
});
await updateTicketRepository(data.ticketId, {
updatedAt: new Date(),
});
await aiQueue.add(
GENERATE_TICKET_SUMMARY_JOB,
{
ticketId: data.ticketId,
messageId: message.id,
},
{
attempts: 3,
backoff: {
type: "exponential",
delay: 5000,
},
removeOnComplete: true,
removeOnFail: false,
},
);
return message;
};
There is an important architectural detail here.
The HTTP request does wait for queue publication, but it does not wait for the AI analysis itself.
Conceptually:
POST message
│
▼
Authorize access
│
▼
Persist message
│
▼
Update ticket timestamp
│
▼
Publish BullMQ job
│
├──────────────► HTTP response
│
▼
Background worker
│
▼
AI processing
That removes LLM execution time from the main request path while still ensuring the API knows whether it successfully handed work to the queue.
It also allows the worker lifecycle to be operated separately from the HTTP server.
What the Background Worker Actually Does
BullMQ uses Redis as the queue infrastructure.
A separate worker consumes AI-processing jobs:
export const aiWorker = new Worker(
AI_QUEUE,
async (job) => {
const startedAt = Date.now();
jobStartTimes.set(String(job.id), startedAt);
const ticketId =
typeof job.data?.ticketId === "string" ? job.data.ticketId : undefined;
workerLogger.info(
{
event: "ai_job_received",
jobId: job.id,
jobName: job.name,
ticketId,
attempt: job.attemptsMade + 1,
},
"Job received",
);
switch (job.name) {
case GENERATE_TICKET_SUMMARY_JOB:
await processTicketSummaryJob(ticketId ?? job.data.ticketId);
break;
default:
workerLogger.warn(
{
event: "ai_job_unknown",
jobId: job.id,
jobName: job.name,
ticketId,
},
"Unknown job",
);
}
},
{
connection: redis as any,
concurrency: 5,
},
);
Inside the ticket-processing workflow, messages are loaded in conversation order and transformed into the conversation passed to the AI layer.
The relevant part of the processing pipeline is:
const conversation = formatConversation(messages);
const aiResult = await generateTicketSummary(
{
conversation,
},
aiContext,
);
const sentimentResult = await generateSentiment(
{
conversation,
},
aiContext,
);
const priorityResult = await detectTicketPriority(
conversation,
aiContext,
);
const sentiment = validateSentiment(
sentimentResult.sentiment || "",
) as any;
const priority = validatePriority(
priorityResult.priority || "",
) as any;
Summary, sentiment and priority are currently generated sequentially inside a job.
The normalized results are then used to update ticket-level AI fields, and an AI interaction record is created for the summary workflow.
This design gives the application retryable asynchronous execution without coupling LLM latency directly to message creation.
But queues introduce their own problems.
I’ll return to those shortly.
Not Every AI Operation Belongs in a Queue
Once background processing existed, it would have been easy to conclude:
AI is slow, therefore every AI operation should be queued.
I don’t think that is the right abstraction.
The correct execution model depends on what the user is doing.
Ticket analysis is enrichment. It can happen asynchronously.
An agent clicking Generate Reply, however, is explicitly waiting for an answer.
For that workflow, the backend generates the suggestion synchronously:
const conversation = formatConversation(messages);
const aiResult = await generateReplySuggestion(conversation, {
ticketId,
requestId: typeof req.id === "string" ? req.id : undefined,
});
res
.status(200)
.json(new ApiResponse("Reply suggestion generated", aiResult));
The important distinction is:
Background analysis
Message → Queue → Worker → AI → Ticket metadata
Interactive assistance
Agent → Request suggestion → AI → Suggested text → Agent
The second flow doesn’t automatically create or send a support message.
That is intentional.
Human-in-the-Loop AI
A generated support reply can be wrong, incomplete or inappropriate for the context.
So the system treats AI output as a suggestion, not an autonomous action.
On the frontend:
<ReplySuggestion
ticketId={data.id}
onUseReply={(content) => {
setComposerDraft(content);
setComposerKey((current) => current + 1);
}}
/>
<MessageComposer
key={composerKey}
ticketId={data.id}
initialContent={composerDraft}
onMessageSent={() =>
setMessageRefreshToken((current) => current + 1)
}
/>
When the user selects Use Reply, the suggestion is moved into the normal message composer.
The agent can then:
Generate
↓
Review
↓
Edit
↓
Send
The AI does not own the final action.
That boundary matters more to me than simply adding another confirmation dialog.
The application architecture itself separates generation from execution.
This is a pattern I would reuse in many AI-assisted workflows where the output affects another person or an important business process.
AI Provider Logic Belongs Behind an Application Boundary
The application currently uses an OpenAI-backed adapter for operations such as:
- summarization,
- sentiment analysis,
- priority detection,
- reply suggestions.
There is also provider orchestration that can use Gemini as a fallback for selected classes of provider, quota, rate-limit, network or timeout failures.
I don’t want ticket or message business logic to know the details of every provider.
Conceptually:
Business workflow
│
▼
AI operation
│
▼
Provider orchestration
│
├── OpenAI
│
└── Gemini fallback
This boundary becomes useful even if the application never changes providers.
It gives one place to reason about:
- prompts,
- provider failures,
- normalization,
- fallback rules,
- model configuration.
Provider abstraction does not mean every model is interchangeable.
Different models behave differently.
The goal is simply to stop provider-specific concerns from leaking through the rest of the application.
Authentication Is Not Authorization
The application uses JWT authentication, but authentication alone does not answer:
Is this user allowed to access this particular ticket?
A valid token identifies the user and role.
Resource authorization still has to be enforced.
For example:
- customers should only access their own tickets,
- agents operate on tickets available to their role and assignment rules,
- admins have broader access.
The frontend also hides or exposes functionality according to role, but that is a UX decision—not a security boundary.
The backend remains authoritative.
This distinction is especially important in React applications because hiding a button is easy to mistake for enforcing permission.
It isn’t.
Redis Has More Than One Responsibility
Redis is not used only because BullMQ needs it.
In this project it also supports:
- short-lived ticket caching,
- distributed rate-limit counters.
That makes Redis shared infrastructure serving different concerns:
Redis
├── BullMQ queue infrastructure
├── ticket cache
└── rate-limit counters
Those responsibilities should still remain conceptually separate even if they use the same underlying technology.
A queue is not a cache.
A cache is not a source of truth.
The database remains the authoritative store for ticket and conversation data.
What the Architecture Gets Right
There are several boundaries in this design that I would keep.
1. User-facing writes are separated from AI enrichment
Message creation does not wait for summary, sentiment and priority generation.
2. Interactive AI and background AI use different execution models
The system chooses between synchronous and asynchronous execution based on the user workflow rather than treating all AI calls identically.
3. Server state remains server state
TanStack Query manages API-backed ticket state rather than copying everything into a global client store.
4. AI suggestions do not automatically become actions
Generated replies remain editable suggestions.
5. HTTP and worker lifecycles are separate
The API server and AI worker can fail, restart and evolve independently at the process level.
But none of those decisions makes the architecture finished.
Several weaknesses become more important as concurrency and operational requirements increase.
What I Would Change Before Scaling
This is the part of the architecture I find most useful to examine.
Queues solve one category of problem while creating others.
1. Fix the Database-to-Queue Consistency Gap
Look again at message creation:
Write message to PostgreSQL
↓
Update ticket
↓
Publish BullMQ job
These are separate operations.
Imagine:
- PostgreSQL successfully saves the message.
- Redis becomes unavailable before
aiQueue.add()succeeds.
The application now has committed business data without the corresponding background event.
Retries do not solve this consistency problem.
A stronger design would introduce a transactional outbox.
Conceptually:
PostgreSQL transaction
┌────────────────────────────┐
│ Save message │
│ Save outbox event │
└──────────────┬─────────────┘
│ commit
▼
Outbox publisher
│
▼
BullMQ
│
▼
Worker
The message and the intent to perform AI processing would be committed atomically in PostgreSQL.
A separate publisher could then reliably move pending outbox events into BullMQ.
That doesn’t make the entire system exactly-once.
It closes a specific and important consistency gap.
2. Retries Need Idempotency
The queue currently configures multiple attempts with exponential backoff.
That improves recoverability from transient failures.
But:
retryable does not automatically mean safe to retry.
Suppose a job:
- calls the AI provider,
- updates the ticket,
- writes an AI interaction,
- fails somewhere around those operations,
- retries.
Without explicit idempotency, repeated attempts can repeat side effects.
A stronger implementation would give each logical processing operation a stable identity and make persisted effects safe to replay.
Possible approaches include:
- deterministic job IDs,
- processing-version records,
- unique database constraints,
- upsert semantics,
- explicit processed-event records.
The exact mechanism depends on the workflow.
The architectural requirement is simpler:
A retry should not accidentally turn one logical operation into multiple business effects.
3. Protect Against Same-Ticket Concurrency
The worker has concurrency greater than one.
That’s useful for processing unrelated tickets.
But it also means two jobs for the same ticket can overlap.
Consider:
Ticket conversation V1
│
└── Job A starts
New message arrives
│
Ticket conversation V2
│
└── Job B starts
Job B finishes first → stores V2 analysis
Job A finishes later → stores older V1 analysis
The final ticket can now contain stale AI metadata.
This is not solved by generic worker concurrency configuration alone.
Before scaling the workflow, I would introduce some combination of:
- per-ticket serialization,
- conversation version numbers,
- compare-before-write logic,
- coalescing pending analysis jobs,
- stale-result rejection.
The important principle is that global concurrency and entity-level ordering are different problems.
4. Give the UI an Explicit AI-Result Delivery Mechanism
Background processing creates another question:
How does the browser learn that AI analysis has finished?
The current architecture does not implement a dedicated real-time result-delivery channel.
At larger scale or with a richer UI, I would make that explicit.
Depending on the product requirements:
Simplest
Polling
↓
More event-driven
Server-Sent Events
↓
Bidirectional realtime needs
WebSockets
I would not automatically choose WebSockets.
If the browser only needs server-to-client status updates, SSE may be enough.
Architecture should follow the interaction requirement, not the popularity of the technology.
5. Tighten Cache Freshness Rules
Caching ticket data can reduce repeated reads, but mutable support conversations create invalidation challenges.
Every path that changes data relevant to a cached ticket needs a clear freshness strategy.
Before relying more heavily on caching, I would explicitly document:
- what is cached,
- the cache key,
- TTL,
- which writes invalidate it,
- whether stale reads are acceptable,
- what happens when invalidation fails.
A short TTL is useful, but TTL alone isn’t a complete consistency model.
6. Improve the AI Context and Evaluation Layer
The current AI pipeline is intentionally straightforward.
There are several improvements I would investigate before treating AI behavior as a mature subsystem:
- explicit structured outputs,
- conversation-size/token management,
- better context construction,
- prompt/model version tracking,
- evaluation datasets,
- regression testing for AI behavior,
- provider health telemetry,
- clearer usage accounting,
- privacy and retention policies for model inputs and outputs.
I would add these based on actual product requirements rather than building an elaborate AI platform prematurely.
The important shift is from:
Does the model return something?
to:
Can I measure whether this AI behavior remains useful and reliable as the application changes?
What I Would Not Add Yet
Scaling discussions can easily turn into architecture shopping lists.
I would not automatically add:
- Kafka,
- Kubernetes,
- multiple microservices,
- vector databases,
- WebSockets,
- event sourcing,
- multiple caches.
None of those technologies is inherently an upgrade.
For this system, I would first strengthen the guarantees around the architecture that already exists:
DB/queue consistency
↓
Idempotent processing
↓
Per-ticket ordering/versioning
↓
Observable job state
↓
AI evaluation
↓
Scale infrastructure when evidence requires it
Complexity should be purchased with a requirement.
The Most Important Lesson: AI Is a Workload, Not the Architecture
It is tempting to describe this application as an OpenAI integration.
That misses most of the engineering.
The LLM is one dependency inside a larger system.
The architecture still has to answer ordinary software-engineering questions:
- Who owns state?
- Which operations block the user?
- What should happen asynchronously?
- Where is authorization enforced?
- What happens after partial failure?
- Are retries safe?
- Can concurrent work overwrite newer state?
- How does the UI learn that background work finished?
- Where does human approval belong?
Those questions existed before LLMs.
AI simply makes several of them more visible because model calls are external, comparatively slow, probabilistic and failure-prone.
Final Architecture
The current design can be summarized as:
React + TypeScript
│
┌────────────────┼────────────────┐
│ │ │
Auth Context TanStack Query Local UI State
│ │ │
└────────────────┼────────────────┘
│
▼
Express REST API
│
Auth / AuthZ / Validation
│
▼
Application Services
│ │
│ │
▼ ▼
PostgreSQL Redis / BullMQ
│
▼
AI Worker
│
┌───────────┼───────────┐
│ │ │
Summary Sentiment Priority
│ │ │
└───────────┼───────────┘
│
AI Providers
Interactive reply path:
Agent → API → AI reply suggestion → editable composer → human sends
It is intentionally simpler than a large distributed system.
And it has known limitations.
I consider both of those facts healthy.
Good architecture is not about pretending the current system can handle every future requirement.
It is about knowing which guarantees the system currently provides, which ones it doesn’t, and where the next architectural pressure points will appear.
Closing
Building this project changed how I think about adding AI to existing applications.
I wouldn’t start by asking:
Where can I call the LLM?
I would start with:
What role should AI have in this workflow?
Then decide:
- synchronous or asynchronous,
- advisory or autonomous,
- user-facing or background,
- retryable or idempotent,
- cached or authoritative,
- human-reviewed or automatically executed.
Once those boundaries are clear, choosing the API is usually the easier part.
In the next article in this architecture series, I’ll go deeper into BullMQ background jobs for AI workloads—especially retries, idempotency and ordering, because adding a queue is only the beginning of making background processing reliable.
About the author
I’m Ram Ji Tripathi, a senior frontend-focused full-stack engineer working across React, TypeScript, Node.js and AI-enabled applications.
I write about frontend architecture, backend systems, AI engineering, system design and the engineering decisions behind building real software.
More engineering notes and projects are available on my portfolio.