Data Scientist
Pune

About Us
Coditation Systems was founded by a serial entrepreneur, and a team of young talented technologists, some of who have grown to spearhead the organization. With its inception in 2016, we became a boutique technology services and solutions firm specializing in Machine Learning & AI, Data Engineering, and Cloud. We have a team of ninja architects, data scientists, data engineers, and software engineers having decades of collective experience of applying emerging technologies to build cutting edge software products.

What are we looking for?
We are looking for a skilled Databricks Developer with hands-on experience in building and optimizing data pipelines on the Databricks platform. The ideal candidate should have strong expertise in big data processing, Spark, and cloud data engineering, with the ability to work in a fast-paced, data-driven environment.

A Day in the Life
  • Build and optimize predictive models for churn forecasting, revenue contraction, and risk analysis at customer-product levels.
  • Develop end-to-end data science libraries and orchestrate pipelines using Kedro, managing data loading, feature engineering, model training, inference, and explainability.
  • Deploy AI/ML solutions on cloud platforms (GCP / AWS) and expose models as APIs.
  • Collaborate with cross-functional teams to translate business requirements into technical solutions.
  • Monitor data quality, detect drifts, and implement data governance frameworks.
  • Perform image classification, text moderation, and other AI/ML-based solutions for client products.
  • Design, develop, and deploy LLM-based AI agents for business intelligence, automated workflows, and decision support.

What you will need
  • 7+ years of experience in Data data science
  • Programming Languages: Python, SQL
  • Frameworks & Libraries: scikit-learn, PyTorch, XGBoost, LightGBM, pandas, NumPy, Polars, LangChain, LlamaIndex, Crew.ai
  • Databases: Postgres (pgvector, tsvector), Redshift, Snowflake (and/or any other vector db)
  • Strong understanding of data science workflows, including data cleaning, feature engineering, model selection, validation, and explainability (e.g., SHAP).
  • Experience in developing chatbots or automated query systems.
  • Experience in image classification and transformer-based NLP models.
  • Good to have: Cloud & Platforms: GCP, API deployment using Flask / FastAPI
  • Good to have: Experience with RAG, LLM prompt optimisation, and GraphLLMs to reduce hallucinations.

Competencies
  • Proven track record of end-to-end ML project delivery, from raw data to production.
  • Strong analytical skills with experience in clustering, forecasting, and predictive modeling.
  • Experience working with subscription and contractual business models.
  • Ability to optimize LLMs and ML models for performance, latency, and scalability.
  • Excellent problem-solving and communication skills.