Machine Learning Engineering

Machine learning engineering combines software engineering, data science, and artificial intelligence to construct computing systems that learn from very large datasets and improve over time. Machine learning engineers build and scale machine learning models and optimize algorithms to efficiently handle very large datasets to solve a wide range of problems. 


Table of Contents

Machine Learning

What is machine learning engineering?

Machine learning engineering is a branch of engineering that implements ongoing data science developments to address complicated or intricate problems with large amounts of disparate data. This work bridges the gap between software engineers who build data platforms and data scientists who focus on developing algorithms and advancing machine learning models.

Machine learning engineers often have specific domain knowledge to understand models and algorithms in addition to strong data and software engineering backgrounds, so they can deploy solutions to real-world operations and retrain models incrementally over time. Through this, machine learning engineers enable computing systems to learn from data, identify patterns, and make decisions with reduced human monitoring. 

What does a machine learning engineer do?

Machine learning engineering is a continuous process. A machine learning engineer builds a solution to meet a goal, typically involving large amounts of (sometimes messy) data. Often, this means automating tedious experimental, monitoring, or analytic activities to reduce the time spent analyzing the data generated by these activities.

When the algorithms don’t perform as expected or the data change, machine learning engineers will adjust the algorithm to improve performance.

Additionally, machine learning engineers often work to train computing systems to provide explanations for their reasoning—an aspect of trustworthy machine learning.

With the right applications, machine learning engineers can do the following:

  • work with computers to separate useful data from noise
  • increase the predictive power and efficiency of computational models
  • aid in decision-making to address important national and global issues. 
Machine learning engineering visualization with laptop
Machine learning engineers develop sophisticated algorithms that harness vast computing power to improve experiments, analysis, and performance. (Photo by Ipopba | iStock.com)

History of machine learning engineering

Researchers, data scientists, and software engineers have been using machine learning engineering to coordinate their algorithms since the beginnings of machine learning in the 1950s. However, the expansive growth of machine learning engineering can be traced back to a few key events and trends in the 2000s:

2006: Geoffrey Hinton and two of his colleagues published a paper outlining a much faster method for training deep neural networks.

Around 2010: GPUs experienced a significant speedup.  

2012: A new computer competition was presented at the IEEE Computer Vision and Pattern Recognition Conference in 2012, called ImageNet. Participants were tasked with developing algorithmic approaches that could sort a database of over a million images.

The breakthrough came when a researcher combined deep neural networks with the much faster GPUs for his entry to ImageNet. Neural networks “learn” by layering hundreds or thousands of algorithms, where the output of one algorithm becomes the input for the next one, making it possible to model the nonlinear and complex relationships in the real world. However, these layered algorithms require vast amounts of computing power, which made using them unworkable until GPUs could be used to process many pieces of data simultaneously.

2017: The majority of teams submitting entries to ImageNet achieved more than 95 percent image recognition accuracy using deep neural networks and the computational power of GPUs. 

Why does machine learning engineering matter?

Machine learning engineering is useful any time a process requires understanding the entire pipeline of incoming data, how the data are interpreted, a systematic way to store results, and the ability to run models in parallel to synthesize information about complex problems. These applications are used to automate the process of interpreting data from scientific experiments, real-time sensors, and environmental monitoring.

Some examples of machine learning engineering include:

  • fraud detection for credit cards and in banking
  • spam filtering in email
  • improving the reliability of the electric grid
  • analyzing medical images acquired by radiography and X-ray scanning: an IBM study exploring the detection efficiency of lymph node cancer cells found that the combined input from a sophisticated AI system and pathologists dropped the AI’s error rate by 7.5 percent, while dropping the human pathologists’ 3.5 percent error rate to a mere 0.5 percent
  • facial recognition: Meta uses machine learning engineering to cross-correlate trillions of images from a billion users to create meaningful predictions
  • recommendations on Netflix, YouTube, and social media platforms like Instagram, Pinterest, and Tik Tok
  • traffic prediction in navigation apps
  • voice recognition programs like Apple’s Siri or Amazon’s Alexa: here, highly specialized algorithms and architecture work together to rapidly ingest, process, and analyze enormous amounts of data.

Purpose-built machine learning engineering applications are continually being developed for environments that require the use of multiple integrated algorithms at the same time. There is a large potential for errors in these complicated environments, and machine learning engineering can help make the transition from proof of concept through development and testing to deployment for real-world users and systems. 

A person looking at a data visualization
Machine learning engineering is useful any time a process requires understanding the entire pipeline of incoming data, how the data is interpreted, a systematic way to store results, and the ability to run models in parallel to synthesize information about complex problems. (Image by NicoElNino | iStock.com)

What are the benefits of machine learning engineering?

Well-implemented machine learning has the potential to improve efficiency by augmenting human intelligence and increasing our effectiveness. One of the largest benefits of machine learning engineering is its ability to automate routine processes:  

  • scanning millions of shipping containers at a port of entry
  • monitoring fluctuations in electricity demands
  • preventing mental fatigue for operators working long shifts
  • helping doctors scan hundreds of images for diagnosis
  • helping scientists analyze massive datasets that would be impossible for a person to interpret in their lifetime.

Many of these monitoring or pattern recognition processes are well suited to computers and machines, which do not fatigue as they perform repetitive simple tasks. As machine learning engineers build algorithms and pipelines to automate systems in real time, they can improve outcomes for both the quality of gathered data and the types of jobs humans will be able to perform, eliminating the need to perform mentally taxing tasks involving consistent repetitions.

What are the limitations of machine learning engineering?

Machine learning relies on pattern recognition to create predictions for the future, but when conditions quickly change, the nuances of human intelligence excel.

An example is Zillow’s failed attempt to use its machine learning engineering tool Zestimate to create a house-flipping business around the time of the COVID-19 pandemic in 2020. It is difficult to assess and predict home values because the price of a house depends on its type and condition, the variations in its physical location, and both the seller’s and buyer’s timing. During the pandemic, changes in human movement combined with supply chain issues, evolving safety requirements, and an unpredictable market crippled Zestimate’s ability to accurately forecast which homes could be moneymakers.

Machine learning engineering applications are only as accurate and complete as the algorithms and data pipelines they rely on.

Machine learning systems require continued involvement from the machine learning engineer to understand and mitigate unintended biases, guide the applications of these systems toward their original intent, and scale these systems as needed. Unintended biases may skew results and outcomes and cause unintended conclusions. A machine learning system that works on an experimental scale may act differently when exposed to larger amounts of data at a faster rate.   

What are the future applications of machine learning engineering?

Machine learning engineering is still in the early stages. In the coming years, researchers expect a proliferation of machine learning engineering in fields outside tech and social media:

  • more efficient drug design and discovery
  • autonomous systems
  • smart homes
  • higher capacity energy storage materials.

As machine learning innovates in an expanding number of disciplines, machine learning engineering will be needed to integrate the efficiency of machine learning into a production scale. In addition to new machine learning approaches developed in tandem with domain experts, the future of machine learning engineering could include new ways of deploying data pipelines and increasing processing speed. Advances in cloud engineering contribute to efforts to automatically deploy data at scale, and developments in heterogeneous compute power push results from a centralized device or location to disparate locations and various types of devices.

How is PNNL using machine learning engineering?

a technician in safety goggles and ear protection works on a server rack

PNNL is leading the next generation of computing for scientific discovery.       
Explore our Computing & AI story

PNNL’s work in machine learning engineering advances contemporary data analytics and artificial intelligence in both science and national security.

Few-shot learning

Our work in few-shot learning quickly builds machine learning models using small amounts of training examples in image, text, audio, and video datasets. This allows researchers, analysts, and decision-makers to glean significantly more information from expensive, time-intensive, or high-risk situations or experiments.

Security

Digital and physical system security benefits from our holistic machine learning engineering approach for research on unexpected system behavior. This includes unexpected behavior from maliciously modified inputs, hardware systems used to train and deploy machine learning algorithms, and the downstream effects of bias in standardized training data or pretrained models.

PNNL data engineers perform mission critical research and development that enables U.S. Customs and Border Protection (CBP) to analyze the vast amounts of data generated by sensors located in thousands of vehicle- and cargo-inspection systems at ports of entry around the country. The cloud-based data pipeline we are developing for the analysis of nonintrusive inspection data from cargo at U.S. border crossings will enable CBP to detect contraband, prevent smuggling, and secure the border.

PNNL has also been a trailblazer in the open-source data analytics space for the better part of the last decade, delivering advanced research and development solutions to sponsors across the national security landscape. These tools and capabilities advance research into understanding deception and misinformation campaigns in social and open-source media. 

Data visualization of a network.
The prevalence of social and open-source data enables the large-scale understanding of topics, sentiments, and social behavior across the globe. PNNL has been a trailblazer in the open-source data analytics space for the better part of the last decade, delivering advanced research and development solutions to sponsors across the national security landscape. (Image by Madelyn Dunning | Pacific Northwest National Laboratory)

Mathematics for AI

Our research in computational topology can be used to build novel mathematical methods for a range of applications, from sensor fusion and anomaly detection to pattern detection and visualization of complex data.

For example, our HyperNetX open-source Python library analyzes and visualizes multiway relationships modeled as hypergraphs. These hypergraphs expose the interconnectedness of the data in areas such as cybersecurity, computational biology, geolocation, and pattern-of-life analysis without artificially generating two-way relationships.

Autonomous research

Materials scientists and data scientists at PNNL are using machine learning engineering to improve the growth of thin films. By integrating machine learning techniques, engineers will be able to detect problems during the growth process faster and correct them on the fly to improve material quality. This real-time feedback is an important part of the development of autonomous experiments.