autodesk

Intern, AI Data Developer (Winter)

autodesk

Toronto · Posted 8h ago

Software Engineering Onsite
Apply now

Location: Toronto, ON, CAN

Job Requisition ID #

26WD101082

Position Overview

As an AI Data Developer Intern on the Advanced Compliance Products (ACP) team at Autodesk, you will help build and scale the data pipelines, models, and platform integrations that power ACP's compliance and analytics capabilities. You will work alongside data engineers, AI/ML engineers, and platform teams to design systems that ingest and process large volumes of telemetry, engineer features for scoring and classification models, and extend shared internal platforms with new domain-specific capabilities; all while ensuring data quality, reliability, and observability throughout the pipeline.

The work we do at Autodesk touches nearly every person on the planet. By creating software for making buildings, machines, and even the latest movies, we influence and empower some of the most creative people in the world to solve problems that matter.

Responsibilities

·      Design, build, and maintain data pipelines (ETL/ELT) that ingest, clean, validate, and transform large-scale telemetry and operational data

·      Engineer features from raw signals to support scoring, classification, or ranking models, and iterate on model performance against benchmark datasets

·      Build and maintain data models, schemas, and versioned datasets that other engineers and analysts can rely on

·      Evaluate model and pipeline output for accuracy, drift, and reliability, and implement monitoring/alerting to catch regressions early

·      Investigate how existing shared platforms and services are structured, identify coupling points and separation-of-concerns gaps, and propose or prototype ways to extend those platforms for new use cases

·      Design, prototype, or build a proof-of-concept extension or integration end-to-end in collaboration with platform stakeholders

·      Document data pipeline architecture, data contracts, and technical decisions to support handoff and future maintainability

·      Present findings, benchmarks, and recommendations to engineering and business stakeholders

Minimum Qualifications

·      Currently enrolled in a full-time undergraduate degree program with expected graduation of April 2027 or later

·      Major in Computer Science, Engineering, Data Science, Statistics, or a related field

·      Proficiency in Python and SQL, including data manipulation and ML libraries (e.g., pandas, Scikit-learn, PySpark)

·      Solid understanding of data engineering fundamentals: ETL/ELT pipeline design, data validation, schema design, and dataset versioning

·      Understanding of Machine Learning lifecycle workflows, model evaluation (precision/recall), and experimental design

·      Experience with Git and familiarity with cloud environments and distributed data processing (e.g., AWS, Azure, Spark)

·      Hands-on experience using AI coding assistants (e.g., Cursor, Claude Code, GitHub Copilot) to write, debug, or refactor code

·      Comfortable working with ambiguous, outcome-based problems rather than a fixed task list

Preferred Qualifications

·      Experience with large-scale/distributed data processing frameworks (PySpark, Spark SQL) against production-scale data

·      Experience with workflow orchestration and pipeline tooling (e.g., Airflow, dbt) and data warehousing concepts

·      Exposure to classification/scoring models, anomaly detection, or feature engineering on structured/behavioral signals

·      Familiarity with GenAI/LLM-powered platforms (e.g., internal chatbots, RAG-based BI tools) and their underlying architecture

·      Understanding of platform/service architecture and separation-of-concerns design in a multi-tenant shared codebase

·      Experience reverse-engineering or “diffing” black-box system outputs against raw inputs to infer underlying logic

·      Practical experience using AI tools (e.g., Cursor, Claude Code) for tasks beyond code generation such as codebase comprehension, test generation, documentation, and data analysis

·      Strong written communication skills for producing architecture docs, data contracts, and stakeholder-facing recommendations

·      Interest in compliance, licensing, or fraud/risk-detection domains

How You'll Use AI to Accelerate This Work

This internship is explicitly outcome-focused, not task-based. You are expected to use AI tools to move faster and go deeper, not just to execute a checklist:

·      Use AI coding assistants (e.g., Cursor, Claude Code, GitHub Copilot) to accelerate pipeline development, feature engineering, and codebase comprehension when working in existing platforms

·      Apply ML methods and AI-assisted analysis to model scoring/classification signals and to diff system outputs against raw inputs, inferring underlying logic

·      Use AI-assisted documentation to produce architecture notes, data contracts, onboarding material, and workflow summaries as you go, rather than writing them from scratch at the end

·      Use AI-assisted analysis to map requirements against existing platform capabilities and quickly surface data quality or architectural gaps

·      Be prepared to explain and justify AI-assisted outputs (model choices, code, documentation) to your mentor and technical stakeholders. AI accelerates the work, but you own the judgment behind it

What You'll Gain

·      Hands-on experience building production-grade data pipelines and models that support real business decisions

·      Exposure to platform architecture, ownership, and contribution models within a large shared internal system

·      A portfolio of measurable outcomes: benchmarked models and pipelines with quantified impact, and a reusable technical playbook

About the Canada Intern Program

The 2027 Canada Internship program runs for 16 weeks (Jan 4 - April 23 2027). All internships are paid. As an intern, you will contribute to meaningful projects, be mentored by industry leaders, and participate in tech talks and other

What they are looking for

Python Sql Pandas Scikit-learn Pyspark

Details

Work type
Onsite

Get new Software Engineering internships by email

Free daily digest, matched to what you pick. Unsubscribe anytime.