We are seeking a hands‑on Data Scientist to develop a matching and recommendation capability that identifies the most likely application associated with a DNS record.
The candidate will combine DNS, CMDB, application, server, IP, ownership, and existing mapped‑record reference data, then apply appropriate data science techniques to generate ranked application matches with explainable confidence scores. The solution should reduce manual investigation and support integration into the broader remediation workflow.
This role supports the Orphan DNS Remediation initiative by applying data science to associate DNS records with the correct enterprise applications.
Key Responsibilities:
- Work with ISRM(Information Security Risk Management) stakeholders and domain specialists to define the DNS-to-application matching problem, business rules, and measurable success criteria.
- Profile, cleanse, normalize, join, and validate data from DNS, CMDB, application, server, IP, ownership, and mapped‑record reference sources.
- Design and compare appropriate matching approaches, including deterministic rules, fuzzy or similarity matching, entity resolution, classification, clustering, and graph‑based analysis where relevant.
- Develop a recommendation engine that returns ranked candidate applications with confidence scores and supporting evidence.
- Validate the solution using confirmed historical mappings and metrics such as precision, recall, top‑k accuracy, coverage, and false‑match rate.
- Perform error analysis and improve features, algorithms, thresholds, and data‑quality rules based on validation results and stakeholder feedback.
- Build reusable Python notebooks and reproducible model pipelines using approved Azure and/or AWS services, with clear technical documentation and knowledge transfer.
Required Qualifications:
- Strong hands‑on experience applying data science, statistics, machine learning, or entity‑resolution techniques to matching, classification, or recommendation problems.
- Expert Python skills, including pandas, NumPy, scikit‑learn, Jupyter notebooks, AWS Glue, and data visualization libraries.
- Strong SQL skills for profiling, transforming, joining, validating, and analyzing data from multiple sources.
- Practical experience with data cleansing, feature engineering, similarity scoring, model selection, validation, error analysis, and explainability.
- Experience creating reproducible analytical or machine learning workflows from exploratory analysis through validated prototype.
- Experience evaluating results using relevant metrics and clearly communicating confidence, assumptions, and limitations.
- Experience with Azure and/or AWS data or machine learning services, Agile delivery, and modern development practices.
- Strong analytical thinking, problem-solving, ownership, collaboration, and communication skills.
Preferred Skills
- Experience with entity resolution, record linkage, fuzzy matching, recommendation systems, or graph analytics.
- Experience with Azure Machine Learning, Amazon SageMaker, AWS data services, or comparable enterprise analytics platforms.
- Experience operationalizing analytical solutions using model versioning, monitoring, retraining, CI/CD, or MLOps practices.
- Familiarity with responsible and explainable AI, data privacy, secure analytics, and governance in a regulated enterprise environment.
#J-18808-Ljbffr
Data Scientist - DNS-to-Application Matching employer: Cognizant
Cognizant is an exceptional employer that fosters a collaborative and innovative work culture, making it an ideal place for professionals seeking to excel in their careers. With a strong emphasis on employee growth and development, you will have access to mentorship opportunities and the chance to lead transformative projects in a dynamic environment. Located in a vibrant area, Cognizant offers unique advantages such as a diverse team and a commitment to work-life balance, ensuring that your contributions are both meaningful and rewarding.