{"id":1539114,"url":"https://alion.io/job/mbrdna-data-scientist-adas-analytics-machine-learning","title":"Data Scientist ADAS Analytics Machine Learning","company":{"id":693735,"name":"MBRDNA","domain":"mbrdna.com","url":"https://alion.io/company/mbrdna","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Lever","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"junior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Jose, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":85000,"max_usd":189000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":188},"experience_years_min":2,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Azure","optional":false},{"name":"Delta Lake","optional":false},{"name":"Embeddings","optional":false},{"name":"FastAPI","optional":false},{"name":"LLM","optional":false},{"name":"Machine Learning","optional":false},{"name":"Multimodal AI","optional":false},{"name":"NumPy","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"Scikit-learn","optional":false},{"name":"Semantic Search","optional":false},{"name":"Semantic Search","optional":false},{"name":"SQL","optional":false},{"name":"A/B Testing","optional":true},{"name":"Anomaly Detection","optional":true},{"name":"AWS","optional":true},{"name":"FAISS","optional":true},{"name":"GCP","optional":true},{"name":"GDPR","optional":true},{"name":"MLFlow","optional":true},{"name":"pgvector","optional":true},{"name":"Pinecone","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Prompt Engineering","optional":true},{"name":"PyTorch","optional":true},{"name":"RAG","optional":true},{"name":"Recommender Systems","optional":true},{"name":"Spark","optional":true},{"name":"Structured Outputs","optional":true},{"name":"TensorFlow","optional":true},{"name":"Time Series Forecasting","optional":true},{"name":"VLM","optional":true},{"name":"Weights & Biases","optional":true}],"status":"live","first_seen_at":"2026-09-16T18:11:10Z","employer_posted_date":"2026-09-16","last_verified_at":"2026-10-01T08:41:09Z","board_verified":true,"closed_at":null,"days_open":14,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":14},"description":"As Mercedes-Benz scales ADAS across production fleets, the US ADAS Data & Forensics team is building applied-AI capabilities that go beyond descriptive KPI reporting. These capabilities include vision-language models that analyze SSR Scene Safety Recording footage, LLM-based event classification and reasoning, embedding-based semantic retrieval over driving scenarios, automated scenario discovery, natural-language event descriptions, and model-based ranking that surfaces the most important events from thousands of daily drives. As a Data Scientist, you will develop and deploy these capabilities in production for two stakeholder groups: management, through fleet-performance summaries, trend forecasts, and AI-generated event narratives; and engineers, through granular model-assisted analysis of ADAS calibration, scenarios, edge cases, and behavioral patterns across road types, weather, firmware, and driver cohorts.\nJob Responsibilities:\nApply vision-language models to SSR video and combine video analysis with structured telemetry to create multimodal event representations.\n\nBuild LLM classification and reasoning pipelines that triage events by severity and root cause and generate human-readable summaries of takeovers, safety events, and deactivations.\n\nDesign embedding pipelines and semantic search for similar-event retrieval, and develop unsupervised clustering methods that discover recurring scenarios and edge-case families at fleet scale.\n\nBuild model-based ranking and scoring systems that reduce manual event triage, develop active-learning loops using engineer feedback, and proactively detect fleet-level anomalies.\n\nQuantify the impact of firmware updates and configuration changes on KPIs, segment driver cohorts, and design A/B and quasi-experimental frameworks.\n\nAnalyze campaign effectiveness, build coverage-optimization models, and develop automated fleet-quality scoring.\n\nDeploy models into PySpark and Delta Lake pipelines and the FastAPI analytics API, build evaluation frameworks for foundation-model outputs, and operate on the Azure data platform, including ADLS, Synapse, and Container Apps.\n\nMinimum Qualifications:\nBachelor's or Master's degree in Data Science, Machine Learning, Statistics, Computer Science, or a related quantitative field. A Master's degree or PhD is beneficial but not required; demonstrated experience carries equal weight.\n2-5 years of experience in data science or applied machine learning.\nDepth in at least one of the following: applied foundation models such as LLMs, VLMs, or embeddings; classical machine learning in production; or statistical experimentation.\nStrong Python skills using NumPy, pandas, and scikit-learn, with an emphasis on clean, testable code, and strong SQL skills.\nExperience with ML model development, including feature engineering, model selection, and evaluation on real data.\nSolid statistics knowledge, including hypothesis testing, regression, and experimental design.\nAbility to communicate findings clearly to engineers and management through reports, presentations, and dashboards.\nAbility to turn ambiguous questions into structured analytical approaches.\nPreferred Qualifications:\nStrongly Preferred\nLLM or VLM application experience, including prompt engineering, structured outputs, and evaluation.\n\nEmbedding models and vector similarity for retrieval or clustering.\n\nPySpark for large-scale processing; candidates with strong pandas experience may ramp up.\n\nTime-series analysis or anomaly detection.\n\nNice to Have\nMultimodal foundation models applied to video or image data.\n\nRAG or vector-database systems such as FAISS, pgvector, or Pinecone.\n\nSpatial or geospatial clustering with DBSCAN or HDBSCAN.\n\nRanking or recommendation systems, active learning, PyTorch, or TensorFlow.\n\nDelta Lake or Parquet; FastAPI or model-serving APIs; MLOps platforms such as MLflow or Weights & Biases.\n\nCloud-platform experience in Azure, AWS, or GCP.\n\nVehicle telemetry or automotive-domain experience; campaign analytics or A/B testing at scale.\n\nData privacy, including CCPA or GDPR, for vehicle data and foundation models.\n\nEnglish required; German is an advantage.","description_format":"text","description_chars":4166,"description_truncated":false,"requirements":{"experience_years_min":2,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[{"language":"English","level":"All levels","optional":false}]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-30T20:15:46Z"}],"liveness":{"score":53,"band":"ok","label":"Likely open","p_open":1,"p_active":0.711,"p_room":0.75,"age_days":14,"expected_fill_days":15,"reasons":["conf:9","win:late","comp:junior"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/mbrdna-data-scientist-adas-analytics-machine-learning","json_url":"https://alion.io/job/mbrdna-data-scientist-adas-analytics-machine-learning.json","meta":{"generated_at":"2026-10-01T10:19:51Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1418,"day_limit":5000,"remaining_today":3582,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}