Bykea is on a mission to enhance the movement of people and parcels across Pakistan, leveraging data science to make transportation smarter, faster, and more efficient. As a Data Scientist at Bykea, you will tackle complex challenges spanning transport, logistics, commerce, economics, and fintech. Your work will involve applying machine learning, deep learning, forecasting, optimization, simulation, geospatial data analysis, and large-scale data processing to solve real business problems, directly impacting products and operations. This role offers significant growth and ownership opportunities within autonomous data science units focused on specific business areas, while also contributing to the company’s broader data science capabilities. You will transform ambiguous business questions into analytical problems, develop solutions, assess their impact, and collaborate closely with Engineering, Product, Marketing, and Operations teams to implement them.
Key Responsibilities
- Build predictive models and algorithms: Develop machine learning and statistical models addressing challenges such as ETA prediction, marketplace churn, demand forecasting, customer behavior, pricing, incentives, and marketplace optimization.
- Solve complex marketplace problems: Work on intelligent allocation, passenger-driver matching, supply-demand positioning, dynamic pricing, routing, and scheduling to enhance marketplace efficiency.
- Work with large-scale data: Extract insights and build models using extensive datasets, distributed computing environments, data warehouses, ETL pipelines, and real-time data streams.
- Apply advanced data science techniques: Utilize machine learning, deep learning, geospatial data mining, forecasting, simulation, optimization, NLP, clustering, and classification to solve practical problems.
- Collaborate cross-functionally: Partner with teams across Engineering, Product, Marketing, Operations, Finance, and Customer Service to define problems, metrics, datasets, run experiments, and translate analysis into business impact.
- Build repeatable data solutions: Develop reliable datasets, modeling pipelines, monitoring systems, and analytical tools to consistently extract value from data.
- Experiment and measure impact: Contribute to A/B testing, experiment design, KPI development, business intelligence, marketing effectiveness, ROI analysis, and other measurement frameworks.
- Own problems end-to-end: Manage projects from understanding business context and data preparation through modeling, validation, implementation, monitoring, and iteration, seeking guidance when necessary.
Required Qualifications
- 2–5 years of professional experience in Data Science, Machine Learning, Data Engineering, Business Intelligence, Data Architecture, Data Modeling, or a related technical field.
- Strong understanding of machine learning, deep learning, data mining, statistics, and algorithmic optimization foundations.
- At least 3 years of hands-on experience with SQL and Python or equivalent, demonstrating strong proficiency.
- Experience building and evaluating machine learning models using frameworks such as scikit-learn or Spark MLlib is a plus.
- Proficient in manipulating, querying, and analyzing large datasets and translating findings into actionable business insights.
- Experience implementing and reporting business KPIs in data warehousing or large-scale data environments.
- Familiarity with distributed data processing or big-data technologies like Spark, Hive, or MapReduce is advantageous.
- Strong quantitative and analytical skills with the ability to decompose complex problems and devise practical solutions.
- Working knowledge of machine learning techniques such as clustering, decision trees, classification, regression, and neural networks, including understanding their appropriate applications and trade-offs.
- Ability to work independently, learn quickly, and communicate technical concepts clearly to non-technical stakeholders.
- Detail-oriented with the ability to manage multiple priorities in a fast-paced environment.
- Competent in using AI and automation tools effectively to enhance productivity in analysis, coding, research, documentation, experimentation, and model development, with at least one year of experience.
- Strong communication skills and a collaborative team player mindset.
Preferred Qualifications and Benefits
- Experience building machine learning models at scale using real-time or streaming data pipelines.
- Familiarity with NoSQL databases, PostGIS, stream processing, or distributed computing platforms.
- Experience working with geospatial data, maps, routing, location intelligence, or marketplace data.
- Knowledge of forecasting, optimization, simulation, recommendation systems, NLP, or experimentation.
- Additional programming experience in R or Scala alongside Python.
- Background in high-volume consumer, marketplace, logistics, mobility, e-commerce, fintech, or similar product domains.
- Proven ability to take machine learning models from experimentation to production.
Bykea’s data science environment includes working with Python, SQL, scikit-learn, Spark MLlib, Hive, NoSQL, distributed computing, real-time data pipelines, data warehousing, machine learning, deep learning, and experimentation tools. While the specific tools may vary depending on the problem, the emphasis is on your ability to select and apply the right approach.
Experience & Qualifications
Candidates should have 2–5 years of relevant professional experience in data science or related fields and hold a Bachelor’s or Master’s degree in Computer Science, Software Engineering, Electrical/Computer Engineering, Operations Research, Mathematics, Statistics, or a related quantitative discipline. Ph.D. candidates or graduates are also encouraged to apply. This position is based onsite in Karachi.
Assessment Process
The selection process includes a hands-on technical assessment focused on real-world data science problems. The evaluation emphasizes problem-solving approach, data reasoning, model selection and explanation, trade-off evaluation, and clear communication of solutions rather than just technical correctness.
Success in this role will be demonstrated by:
- Delivering machine learning and analytical solutions that address meaningful business challenges and improve measurable metrics.
- Developing reliable, scalable, and practical models and data pipelines for production use.
- Creating clear datasets, KPIs, experiments, and monitoring systems that enable better decision-making.
- Collaborating effectively with cross-functional teams across the organization.
- Taking increasing ownership of complex data science problems from definition through implementation and impact measurement.
- Continuously improving techniques, processes, and technologies to maximize value extraction from data.