Automated Job Searching & Resume Tailoring System
From a repetitive workflow to an AI application.
A backend-focused AI system that automates job discovery, company research, and resume tailoring — turning a repetitive workflow into a controllable application.
- -60%
- inference time / message
- -30%
- cost / message
- 7
- specialized stages
Focus
Architecture · Backend · Agent workflow · Data model · Automation
Stack
FastAPI · PostgreSQL · OpenAI · Multi-agent · Apify
Context
Applying for jobs looks simple from the outside.
- Find a position.
- Understand the company.
- Tailor the resume.
- Write the application.
But doing this repeatedly turns into a long chain of small, disconnected tasks. The problem wasn't simply “how can an LLM write a resume?”
The real problem was:
How do we turn this entire workflow into a system that can reason through multiple steps and still produce something a user can review and control?
Architecture
The turning point
A single AI call could generate a resume.
But a single prompt couldn't reliably handle job discovery, company research, resume tailoring, criticism, scoring, and proposal writing at the same time.
So instead of building around one large prompt, I treated the workflow itself as the architecture.
- 01Job Discovery
- 02Company Research
- 03Candidate / Job Analysis
- 04Resume Tailoring
- 05Critic & Review
- 06Scoring
- 07Application Proposal
Each stage has a different responsibility and can evolve independently.
The backend became the orchestration layer connecting these stages rather than simply exposing an LLM API.
Engineering decisions
The engineering decisions
Building the system wasn't just about connecting an LLM to a job API. As the workflow grew, several engineering problems started to appear.
01
Don't let one agent own the entire workflow
- Problem
The workflow involved job discovery, company research, resume analysis, tailoring, reviewing, scoring, and proposal writing.
Putting all of these responsibilities into one agent would make the agent's context increasingly large and its behavior harder to control.
- Why
I separated the workflow into specialized agents, each responsible for a specific task.
This made responsibilities explicit and allowed individual stages to have their own context, prompts, and evaluation criteria.
- Trade-off
The architecture became more complex. More agents meant more orchestration, more LLM calls, and more places where failures could occur.
In return, the workflow became easier to reason about, debug, evaluate, and extend.
02
Persist company research
- Problem
Company research can be expensive and repetitive.
If every resume-tailoring workflow independently researched the same company, the system would repeatedly spend time and tokens collecting information that had already been discovered.
- Why
I treated company research as reusable application data rather than temporary context for a single AI run.
Research results could be persisted and reused by later workflows targeting the same company.
- Trade-off
Persisting research introduces data-modeling and freshness problems.
The system now needs to decide when existing research is still useful and when it should be refreshed.
The benefit is that expensive research becomes a reusable asset instead of being regenerated for every application.
03
Separate AI execution from application state
- Problem
LLM calls and other AI operations are long-running, failure-prone, and dependent on external services.
Application state — jobs, companies, resumes, applications, and workflow status — needs to remain consistent even when an AI execution fails or is retried.
- Why
I separated the persistent application state from the execution of AI workflows.
The application owns the state and workflow status, while AI execution operates as a processing layer that reads the required context and produces outputs.
- Trade-off
This separation requires more explicit state management.
The system needs to track execution status, retries, intermediate results, and failures instead of relying on a single request-response flow.
In return, AI jobs can be retried or changed independently without coupling the application's core state to a particular LLM execution.
04
Track token usage, latency, and cost
- Problem
Once a workflow contains multiple AI calls, it becomes difficult to understand why a single application is slow or expensive.
A successful output alone doesn't tell me whether the workflow is efficient.
- Why
I tracked metrics such as token usage, latency, and estimated cost across AI executions.
This gave me visibility into which stages consumed the most resources and provided data for later optimization.
- Trade-off
Instrumentation adds implementation effort and produces additional data that needs to be stored and interpreted.
However, without these measurements, optimization would largely depend on intuition.
For an AI application, execution metrics become part of understanding the system itself.
05
Parallelize independent work
- Problem
Not every step in the workflow depends on the previous one.
Several pieces of company or job information can sometimes be collected independently. Running each operation sequentially would make the total workflow unnecessarily slow.
- Why
I parallelized independent operations where their inputs did not depend on one another. Instead of:
A → B → C → D
some parts of the workflow could become:
┌→ B ─┐ A ────┼→ C ─┼→ D └→ E ─┘This reduced unnecessary waiting between independent operations.
- Trade-off
Parallel execution increases concurrency and therefore increases the number of things that can fail at the same time.
It also requires more careful handling of rate limits, retries, resource usage, and result synchronization.
The trade-off was worthwhile for independent, latency-sensitive operations where sequential execution provided no additional correctness.
06
Deduplicate external job data
- Problem
Job data came from external sources, and the same position could appear multiple times across fetches or sources.
Without deduplication, the system could create duplicate records, repeat downstream processing, and waste AI calls.
- Why
I introduced deduplication at the data-ingestion layer so that the rest of the workflow could operate on a cleaner set of jobs.
This prevented duplicate data from propagating into company research, analysis, and resume-tailoring stages.
- Trade-off
Deduplication requires defining what actually makes two job records the “same” job.
A rule that is too strict can leave duplicates; a rule that is too aggressive can incorrectly merge different positions.
The system therefore needs a stable identity strategy and a balance between deduplication accuracy and implementation complexity.
Outcome
What I built
- FastAPI backend
- PostgreSQL data layer
- Job-search API integration
- Company research workflow
- Multi-agent orchestration
- Resume tailoring
- Critic / reviewer / scoring stages
- Application proposal generation
The result was not another LLM demo.
It was a backend application that connected external data, research, multiple AI stages, persistent state, and generated outputs into one workflow.
When AI becomes part of a product, the difficult part is often not calling the model. It is designing the system around the model.
Engineering focusArchitecture & workflow design
Read the build logs for this projectNext · 02 · Deploy AI systems
Lecture Video Generation Application