Data Scientist (SP6 - SP10)
Position summary
Introduction
Job description
KEY PERFORMANCE AREAS (KPAs)
| Data Scientist Practitioner | SENIOR Data Scientist | LEAD Data Scientist |
1. Work Complexity | · Designs experiments, test hypotheses, and build models. · Conducts data analysis and develops moderately complex models using Python (scikit-learn, pandas) and PySpark on the Azure/Fabric platform. | · Designs experiments, test hypotheses, and build models. · Conducts advanced data analysis and complex designs algorithm. | · Designs experiments, test hypotheses, and build models. · Conducts advanced data analysis and highly complex designs algorithm. · Applies advanced statistical and predictive modelling techniques to build, maintain, and improve on multiple real-time decision systems.
|
2. BUSINESS REQUIREMENTS | · Works with stakeholders to identify the business requirements and the expected outcome. · Works with and alongside business analysts by suggesting other products of interest to the client. · Models and frames business scenarios that are meaningful and which impact on critical business processes and/or decisions. | · Works with stakeholders to identify the business requirements and the expected outcome. · Works with and alongside business analysts by suggesting other products of interest to the client. · Models and frames business scenarios that are meaningful and which impact on critical business processes and/or decisions. | · Leads discovery processes with stakeholders to identify the business requirements and the expected outcome. · Works with and alongside business analysts by suggesting other products of interest to the client. · Models and frames business scenarios that are meaningful and which impact on critical business processes and/or decisions.
|
3. DATA REQUIREMENTS | · Collaborates with subject matter experts to select the relevant sources of information. | · Identifies what data is available and relevant, including internal and external data sources, leveraging new data collection processes such as smart devices and geo-location information or social media. · Collaborates with subject matter experts to select the relevant sources of information. · Works with IT teams to support data collection, integration, and retention requirements based on the input collected with the business. | · Identifies what data is available and relevant, including internal and external data sources, leveraging new data collection processes such as smart devices and geo-location information or social media. · Collaborates with subject matter experts to select the relevant sources of information. · Makes strategic recommendations on data collection, integration and retention requirements incorporating business requirements and knowledge of best practices.
|
4. ANALYSIS | · Works with colleagues to solve client analytics problems and documents results and methodologies. · Works in iterative processes within a team and validates findings. · Performs experimental design approaches to validate finding or test hypotheses. · Validates analysis by comparing appropriate samples. · Employs the appropriate technique to discover patterns — traditional ML (e.g. XGBoost/LightGBM) for structured data and AI methods (RAG, prompt engineering, Agent orchestration, etc.) for structured and unstructured data. · | · Solves client analytics problems and communicates results and methodologies. · Works in iterative processes with the client and validates findings. · Develops experimental design approaches to validate finding or test hypotheses. · Validates analysis by comparing appropriate samples. · · Employs the appropriate technique to discover patterns — traditional ML (e.g. XGBoost/LightGBM) for structured data and AI methods (RAG, prompt engineering, Agent orchestration, etc.) for structured and unstructured data. | · Develops innovative and effective approaches to solve client's analytics problems and communicates results and methodologies. · Works in iterative processes with the client and validates findings. · Develops experimental design approaches to validate finding or test hypotheses. · Validates analysis using scenario modelling. · Identifies/creates the appropriate technique to discover patterns — traditional ML (e.g. XGBoost/LightGBM) for structured data and AI methods (RAG, prompt engineering, Agent orchestration, etc.) for structured and unstructured data. |
5. Qualification and Assurance (Data Quality) | · Uses the expected qualification and assurance of the information to quantify the accuracy metrics of the analysis. | · Assesses, with the business, the expected qualification and assurance of the information in support of the use case. · Defines the validity of the information, how long the information is meaningful, and what other information it is related to. | · Assesses, with the business, opportunities to enhance the qualification and assurance of the information to strengthen the use case. · Defines the validity of the information, how long the information is meaningful, and what other information it is related to. |
6. ACCESS MANAGEMENT AND CONTROL | · Qualifies where information can be stored or what information, external to the organisation, may be used in support of the use case. | · Works with the data steward to ensure that the information used is in compliance with the regulatory and security policies in place. · Qualifies where information can be stored or what information, external to the organisation, may be used in support of the use case. | · Works with the data steward to ensure that the information used is in compliance with the regulatory and security policies in place. · Qualifies where information can be stored or what information, external to the organisation, may be used in support of the use case. |
Minimum requirements
| Data Scientist Practitioner | SENIOR Data Scientist | LEAD Data Scientist |
QUALIFICATIONS & EXPERIENCE | · Bachelor’s degree in mathematics, statistics or computer science or related field. · Typically requires 1-3 years’ experience manipulating large datasets and using databases · 1-3 years' experience in Python (scikit-learn, pandas, etc.) and SQL, with exposure to PySpark. Experience with Azure Machine Learning, MLflow and feature stores. · Experience in statistical analysis, quantitative analytics, forecasting/predictive analytics, multivariate testing, and optimization algorithms. · Familiarity with cloud data platforms — Microsoft Azure, Fabric/OneLake and the medallion (Bronze/Silver/Gold) lakehouse — and distributed processing with Spark. · Demonstrable ability to quickly understand new concepts-all the way down to the theorems- and to come out with original solutions to mathematical issues. · Good communication and interpersonal skills. · Knowledge of one or more business/functional areas. · Relevant Microsoft certifications advantageous: Azure Data Scientist, AI Engineer, Fabric
| · Bachelor degree in mathematics, statistics or computer science or related field; Master degree preferred. · Typically requires 3-5 years of relevant quantitative and qualitative research and analytics experience. · Solid knowledge of statistical techniques. · The ability to come up with solutions to loosely defined business problems by leveraging pattern detection over potentially large datasets. · Strong programming skills in Python and PySpark, with hands-on Azure Machine Learning, MLflow and CI/CD (Azure DevOps); experience across traditional ML and AI methods. · Proficiency in statistical analysis, quantitative analytics, forecasting/predictive analytics, multivariate testing, and optimization algorithms. · Strong communication and interpersonal skills. · Knowledge of one or more business/functional areas. · Relevant Microsoft certifications advantageous: Azure Data Scientist, AI Engineer, Fabric | · Masters in mathematics, statistics or computer science or related field; PhD degree preferred. · Typically requires 5 or more years of relevant quantitative and qualitative research and analytics experience. · Solid knowledge of statistical techniques. · The ability to come up with solutions to loosely defined business problems by leveraging pattern detection over potentially large datasets.Expert programming in Python/PySpark; deep experience deploying and governing ML and AI/agentic solutions in production (Azure ML, MLflow, Foundry). · · Proficiency in statistical analysis, quantitative analytics, forecasting/predictive analytics, multivariate testing, and optimization algorithms. · Strong communication and interpersonal skills. · Experience leading teams. · In-depth industry/business knowledge. · Relevant Microsoft certifications advantageous: Azure Data Scientist, AI Engineer, Fabric and a Foundry/GenAI pathway for senior levels. |
