Beyond answers: New Genie One features to turn insights into action
We all know the pattern: you ask an AI tool a question and get an answer in seconds,...
Tracking the latest research and engineering blogs in the Data Engineering space.
We all know the pattern: you ask an AI tool a question and get an answer in seconds,...
The August 2026 release of the on-premises data gateway is version 3000.330. The new August 2026 release of the on-premises data gateway (version 3000.330) c...
Construction generates abundant data, from equipment telemetry and maintenance records...
When data grows faster than the systems around it, complexity becomes the default....
By Emma Yanyang Kong, Aditya Deshpande, Asad Abbasi, Bowei Yan, David Fagnan, Ashish Rastogi, Dhaval Patel, Ray Zhang Introduction The Netflix experience is ...
RDBMS (Relational Database Management System) databases face several limitations, including slow execution with multi-hop queries and a lack of explainabilit...
Deploying natural-language interfaces over enterprise OLTP catalogs fails at scale because semantic parsers collapse under schema-graph scaling, inflating co...
In the realm of Explainable AI, classification results are often explained via counterfactuals (CFs for short), which are (ideally small) perturbations to an...
Efficiency of skyline algorithms is highly influenced by the underlying data characteristics. Traditionally, optimization efforts have focused on minimizing ...
Set reconciliation recovers the symmetric difference $A\triangle B$ with communication far below the data volume. Invertible Bloom Lookup Tables (IBLTs) are ...
We study size bounds for conjunctive query (CQ) results which in recent years have played a crucial role in database theory. In particular, we compare the so...
Continuous aggregate queries over sliding windows are common in real-time analytics, but most systems report \emph{what} an aggregate is doing without attrib...
An LLM call in a semantic data processing system is expensive enough to dominate query cost, yet slow enough to hide a CPU-side learner's update behind its r...
Data verification, the process of labeling data items as correct or incorrect, is a preprocessing step that may critically affect the quality of results in d...
DeQL (Decision Query Language) extends SQL to express decision queries: given options drawn from relational data, constraints from policy, and a measurable o...
Configuration tuning is critical to database performance but remains difficult in real deployments. Despite notable advances, prior methods still leave subst...
Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation f...
Multi-vector retrieval has become an important primitive for fine-grained matching in information retrieval, with emerging applications in areas such as reco...
GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks...
Large Language Models (LLMs) are emerging as promising assistants in High-Performance Computing (HPC), where programming remains complex and expertise-intens...
Extracting a shared set of unknown, not directly measurable quantities from multiple, heterogeneous datasets is a common challenge across scientific domains....
Chunked prefill pipeline parallelism (CPP) is a key technique for LLM inference. However, equal-size chunks exhibit imbalanced latency, as later chunks atten...
The use of API gateways within geographically distributed multi-cloud Kubernetes clusters poses a tradeoff between infrastructure cost, computational resourc...
This paper measures the performance impact of running large language model inference and training inside a Trusted Execution Environment (TEE) on NVIDIA B200...
Byzantine fault-tolerant (BFT) systems are, in principle, an appealing foundation for transactional applications involving mutually distrustful participants....
Whole-slide images (WSIs) are central to computational pathology but are prohibitively large, making patch-based processing the practical unit for foundation...
We study binary consensus in the \emph{stochastic broadcast model}, which assumes $n\geq 2$ processes communicating synchronously by message broadcasts. At e...
Many irregular neural workloads induce skewed many-to-one reductions with repeated neighborhoods and nonlocal communication. Conventional NoC mappers optimiz...
Autonomous agents increasingly perform bounded software tasks under an orchestrator that retries, resumes, and budgets them. The machinery such orchestrators...
Distributed training across a wide area network (WAN) is challenging, as continuous parameter exchange by islands of compute is constrained by limited bandwi...
Set reconciliation recovers the symmetric difference $A\triangle B$ with communication far below the data volume. Invertible Bloom Lookup Tables (IBLTs) are ...
Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) wit...
Distributed Quantum Computing (DQC) addresses the physical scaling limitations of monolithic quantum processors by networking modular Quantum Processing Unit...
Reliable mobile GUI agents must retain and reuse information across actions, applications, and repeated interactions. However, current benchmarks systematica...
We show that almost stable matching can be solved in constant distributed rounds on general bipartite graphs $G=(V,E)$ using only a few shared random bits. S...
Tensor decomposition (TD) is essential for analyzing high-dimensional sparse data, yet its irregular computations and memory-access patterns pose major perfo...
Large-scale machine learning workloads increasingly rely on multi-GPU systems, yet their performance is often limited by an overlooked component: the CPU. Th...
Euler-Lagrange (EL) simulations provide a direct and robust framework for modeling disperse multiphase flows. However, they are computationally expensive. Wh...
Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities...
Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation f...
At scale, your training efficiency is determined by a single metric: "goodput", the...
Welcome to the August 2026 Fabric update! Microsoft Fabric continues to evolve with new capabilities that help organizations build, manage, and scale their d...
Welcome to the fifth post in our Business Events, Fabric Events, and Azure Events series for Microsoft Fabric. This series takes you from foundational event-...
We are headed to VLDB 2026 to share multiple innovations that power the Databricks platform...
Limited-time offers can be a powerful way to create excitement, bring guests back,...
The MotivationMore and more enterprises are now asking agents to work with their...
In the first blog of this series, we looked at how Lakebase Postgres is rewriting...
Agents that interact with a traditional OLTP database often create bottlenecks at the storage layer. New deployments...
IDENTITY columns in Fabric Data Warehouse are now generally available. Since preview, thousands of customers have adopted IDENTITY columns to simplify data w...
Reads travel from a shared model to agents, apps, and people. Writes enter through an audited action gate and write back to the systems of record that own th...
Your FinOps lead needs to easily drill-down into Databricks spend and identify what’s driving costs, ...
Twenty sessions. Four days. One short link! The Microsoft SQL product team is heading to Barcelona for SQLCon Europe 2026, September 28 to October 1, co-loca...
Fabric Apps now support anonymous data access to enable public-facing experiences without mandatory sign-in. This feature balances accessibility with securit...
One of the most common questions I hear from customers as they move to an allocation-based billing model is:"What levers do I actually have if I want to cont...
Microsoft OneLake, the single, unified, logical lake for all of your organization’s data, can now work with more data sources! We’re excited to announce the ...
How to Build a Data Platform From Scratch We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams operating under a single p...
Samuel Yeboah, Francesco Di Chiara and Mingliang Liu Today, Netflix runs two Flink autoscalers. That is exactly one more than we want. We built the first one...
Connecting streaming workloads to Azure Event Hubs should not require teams to manage shared credentials. Organizations building real-time data solutions nee...
Operational dashboards often need to answer a direct question: Is this metric healthy, approaching a limit, or already in a critical range? A number alone sh...
AI agents are increasingly becoming part of how developers work with data. But generating SQL is only one step in a typical warehousing workflow. Users still...
Part two of a series on medallion architecture with Fabric Data Warehouse One job per layer — with enough implementation detail to make it real In Part 1 of ...
Enterprise Planning is cyclical, and Fabric Planning is built for that reality. Instead of per-user licenses or a separate subscription, Fabric Plan uses an ...
Microsoft Fabric encrypts all data at rest by default with Microsoft-managed keys. For organizations with strict compliance and regulatory requirements, Cust...
This blog post focuses on connectivity options for Microsoft Fabric workloads that use Data Factory runtime components, including Dataflow Gen2, Fabric Data ...
Microsoft Fabric & Microsoft SQL Community Conference, the premier gathering for data professionals, developers, and business leaders who are shaping the fut...
The Fabric data agent answers natural-language questions about your data and returns visuals, such as charts and graphs, alongside its text responses. Those ...
When a real-time data pipeline breaks, the first question is always the same: what went wrong, and where? Eventstream observability in Microsoft Fabric bring...
How your data agent decides where to look
Query Acceleration for Microsoft Fabric Data Warehouse is a new GPU-powered capability that helps accelerate analytical workloads automatically. Query Accele...
How to Build a Data Platform From Scratch We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams operating under a single p...
The word ontology now covers almost every kind of semantic ambition in data and AI. It can mean a metrics layer that consistently calculates “revenue”. It ca...
How to Build a Data Platform From Scratch We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams operating under a single p...
How and Why Netflix Built a Real-Time Distributed Graph: Part 3 — Querying the graph with gRPC execution API Authors: Nilesh Mishra and Ajit Koti This is the...
https://www.reddit.com/r/LinusTechTips/comments/13qji93/user_benchmark_meme/ A benchmark is useful only when you can explain what it measured, why it stopped...
How to Build a Data Platform From Scratch We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams operating under a single p...
by Aarti Laddha, Richard Diaz-Cool, Rishika Idnani, Venkatesh Selveraj Netflix supports a vast and evolving set of features and content types, ranging from 4...
Authors: Ying Li, Arjun Rao, Shradha Sehgal Introduction Recommendations sit at the heart of the Netflix experience. Our current production models rely on th...
Editor’s Note: leetdata.ai & aidataengineer.io updates Last week we were busy fixing some bugs and upgrading a few features. aidataengineer.io now supports s...
How to Build a Data Platform From Scratch We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams, operating under a single ...
By AI Platform’s Model Runtime team and Inference team Introduction Most organizations consume LLMs through hosted APIs. Netflix went further — we run the fu...
By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva & Nathan Fisher A deep dive into the engineering challenges of building a real-time service d...
Deep Dive: See how Dagster uses AI internally What does it actually look like to use AI as part of your day-to-day engineering workflow? At Dagster, we’ve de...
The hard problem is not making personal data disappear. It is preserving the minimum useful properties for an allowed purpose, limiting who can link the data...
Deep Dive: See how Dagster uses AI internally What does it actually look like to use AI as part of your day-to-day engineering workflow? At Dagster, we've de...
A software engineer preparing for a career walks a well-worn path: grind LeetCode, study system design primers, rehearse mock interviews on platforms built f...
Authors: Lequn Wang, Jiangwei Pan, and Linas Baltrunas Figure 1. Autoregressive homepage generation. GenPage builds a Netflix homepage one row or entity at a...
How to Build a Data Platform From Scratch We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams, operating under a single ...
Two years ago, Maya finished a twelve-week data engineering bootcamp. She crushed it. She built a batch pipeline, wired up Airflow, wrote her first dbt model...
By Zhuoning Yuan, Ta-Ying Cheng, Benjamin Klein, Bahareh Azarnoush Introduction At Netflix, we build technology to help storytellers bring their creative vis...
By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha Chandrashekar As a part of the journey to transition Netflix’s compute infrastructure to b...
How to Build a Data Platform We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams, operating under a single platform. In ...
How to Build a Data Platform We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams, operating under a single platform. In ...
How to Build a Data Platform We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams, operating under a single platform. In ...
How to Build a Data Platform We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams, operating under a single platform. In ...
How to Build a Data Platform We wrote an eBook on Data Platform Fundamentals to help you be like the happy data teams, operating undering a single platform. ...