Wiki
Job Descriptions
Reading and writing data job descriptions: role clarity, problem framing, requirements, red flags, and candidate fit.
Related Wiki Pages
A job description is the shared role spec between a hiring team and a candidate. It names the business problem and team context. It also names the level, responsibilities, required evidence, and interview path. In data work, the description also separates noisy titles such as Data Scientist, Data Engineer, and Data Analyst from the actual work.
The posting works as both a hiring spec and a candidate diagnostic. Data science recruiters build the spec with hiring managers. They then adjust requirements against market reality.[1] Candidates can treat a mismatch between title, responsibilities, and team context as a warning. The company may not have named the data problem it needs to solve.[2]
Role Clarity
A useful data job description names the team, the business problem, and the level. It also names the must-have work before it lists tools. Recruiters and hiring managers define what the role needs, then check whether the market can supply those requirements.[1]
Posting language should focus on the problems the person will solve. It should not stop at the perks the company can offer.[1] That places job descriptions inside Hiring and CV Screening.
Candidates read the same document from the other side. In a targeted search, they start with the company problem and map their evidence to it. They favor fewer relevant applications over broad application volume.[3] The interview version is similar. The posting helps candidates decide which CV points and examples belong in the conversation. It also helps them choose questions.[4]
Candidates can use a posting to infer whether a company wants product data science or machine learning engineering. The posting can also reveal analytics work or another role structure. [4] Keyword-driven recruiting can miss strong people from the candidate side. That happens when screening rewards tool matches instead of the actual problem and role fit. [5]
The same keyword trap can swing the market from one noisy label to another. A company can overcorrect from “data scientist” to “data engineer”. It can still miss the capability it needs if the posting never names the work or team boundary. [6].
The ABC framework gives hiring teams a way to avoid that trap. If the role is Analyst-shaped, ask for exploration, visualization, and storytelling evidence. If it’s Builder-shaped, ask for production ownership, MLOps practice, and cloud delivery. If it’s Consultant-shaped, ask for stakeholder persuasion, business framing, and leadership examples [7]. Those requirements describe the work behind the title, so candidates can decide which evidence to show.
Hiring teams should be just as specific with newer client-facing titles. A forward deployed engineer posting shouldn’t read like a generic AI engineer or machine learning engineer role. Hiring managers should name customer deployment work, product adaptation, and the expectation that repeated client needs become reusable product enablers [8].
Requirements and Level
Requirements should describe the work before the technology stack. For data scientists, the posting should distinguish experiments and modeling from product decision support and deployed ML systems. For analysts, it should separate BI reporting and product analytics from stakeholder analysis and Analytics Engineering.
For teams trying to hire data engineers, the description should name the operating surface, such as pipelines and platform infrastructure. It may also include data models, governance controls, or production ownership.
Level matters as much as title. Junior data-engineering descriptions should leave room for training and mentorship. Senior descriptions can ask for system ownership, architecture judgment, and cross-team communication.[9]
Titles are noisy, so hiring teams can check SQL and Python first. They can then weigh problem framing, outcomes, projects, and level-appropriate responsibility above the label alone.[9]
A posting that asks for junior compensation with senior scope damages both sides of the market. Candidates reading Job Search signals may self-select out or tailor the wrong evidence. Recruiters applying CV Screening criteria may also reject people against a role that was never clearly defined.
Candidate Evidence
A strong description points candidates to the evidence recruiters screen for. That evidence includes experience, education, clear responsibility, and a readable CV.[1]
Candidates can use those signals only when the hiring team states what the person will actually do.
For experimentation roles, the description should name experiment design and metrics. It should also name product decision work. For engineering roles, it should name pipeline work and data quality. It should also state ownership boundaries and orchestration expectations.
Tools such as SQL and Python can then act as evidence for a concrete job. Airflow, dbt, cloud platforms, or vector databases can do the same when the hiring team links each tool to the work behind it. They shouldn’t appear as a loose keyword list.
Role requirements should leave room for valuable non-CS evidence when the work benefits from it. A sociology background or qualitative interviewing practice can strengthen data science work. Domain practice can do the same when the person also has the needed statistics and programming base [10]. Job descriptions that only scan for degree names or tool strings can miss that fit.
Candidates should check industry fit and use-case alignment, then show projects with business impact.[3]
The CV then works like a landing page for the role rather than a complete career inventory.[4]
The job description gives enough signal to put relevant achievements first and remove unrelated detail.
Role-Mismatch Signals
A misleading title can hide a different role.[2] A “data scientist” posting dominated by ETL or platform work may be Data Engineering. A “data analyst” posting that owns instrumentation and dbt models may be closer to Analytics Engineering. Semantic layers point in the same direction. The same ambiguity appears in broader Data Teams questions when candidates can’t tell whether analysts, engineers, ML engineers, and product stakeholders already exist.
Long technology lists and vague responsibilities are another warning sign.[2]
Tools matter, but the description should explain why they matter. Airflow often means batch pipelines. dbt often means analytics engineering. Vector databases may mean search or retrieval-augmented generation. Without that context, the tool list becomes keyword noise.
Hiring teams also reveal role design through language. Inclusive wording affects who sees themselves in the role.[1] Olga Ivina gives the data-science hiring version. Teams can attract more diverse candidate pools by reviewing wording and requirements. They should remove discouraging phrases before the post reaches the market [11] [12]. Words such as “rockstar” and “ninja” can signal unclear expectations. They can also signal hero culture or a narrow view of who belongs in the role.[2]
Compensation and Interview Context
Descriptions are weaker when they hide salary and interview context. Candidates need fit information before recruiter calls. Salary transparency and salary bands belong in the fit conversation. Negotiation, leveling, and role breadth belong there too.[2][1] Use Salary Negotiation for the same salary-range and leveling questions.
Interview structure is also part of the role signal. Candidates should see recruiter screening and interview rounds early enough to judge the time investment.[4]
Take-home workload belongs in that same signal because it determines candidate time investment.[4]
Data-engineering assessments should also vary by level.[9]
A good posting tells candidates enough about the interview path to judge whether the requested work is proportionate to the role.
Portfolio Fit
Portfolio evidence should answer the job description, not display unrelated work. Projects are stronger when they connect to concrete use cases and show business impact.[3]
Candidates without direct industry experience can use cold-start projects. Synthetic data and blogging can provide role-relevant evidence.[4]
Shareable data-engineering projects and GitHub work can show pipeline thinking, privacy awareness, and clear storytelling.[9]
For Data Engineering Portfolio Projects, strong evidence includes ingestion and orchestration. Tests, data quality checks, and a runbook make the project easier to evaluate.
For Machine Learning Portfolio Projects, the evidence should show model framing and evaluation. Deployment, monitoring, and tradeoffs make the work closer to a real role.
For Analytics Engineering Portfolio Projects, the strongest evidence is clean models and metric definitions. Stakeholder-facing documentation and decision support show how the work would be used.