Business AI Research
Featured research
Foundation Models for Tabular Data within Systemic Contexts Need Grounding
We propose Foundation Models for Semantically Linked Tables (FMSLT) to advance the understanding of structured enterprise data. Enterprise tables are interconnected through operational logic and semantic relationships that define how businesses operate. Recognizing and modeling these connections is essential for capturing the true nature of enterprise data.
ConTextTab: A Semantics-Aware Tabular In-Context Learner. NeurIPS 2025 (Spotlight)
By leveraging LLM embeddings, ConTextTab successfully integrates semantics from table features into tabular prediction, shining on data with high semantic content such as free text or descriptive categories.
RELATE: A Schema-Agnostic Perceiver Encoder for Multimodal Relational Graphs
We introduce RELATE (Relational Encoder for Latent Aggregation of Typed Entities), a schema-agnostic, plug-and-play feature encoder that can be used with any general purpose GNN.
Open-Source Enterprise Datasets
We introduce SALT and SALT-KG, the first enterprise datasets built from real customer ERP systems. They combine rich, linked business tables with a curated knowledge graph capturing semantic context. Together, they lay the foundation for advancing foundation models that truly understand structured enterprise data.
Paper and Open Source Foundation Model on Tabular Data: SAP-RPT-1-OSS
We have published our ConTextTab research paper at NeurIPS 2025 (spotlight paper) and provided an open weight version of our model as SAP-RPT-1-OSS.
Who we are
At SAP Business AI Research, we serve as the bridge between academia and industry, dedicated to advancing next-generation AI systems. Our research addresses the complexities of real-world enterprise environments by integrating cutting-edge AI techniques with domain-specific challenges. We focus on two main research tracks to ensure that our models are not only powerful but also practical, trustworthy, and scalable.
Research areas
Track A: Structure - Aware Foundation Models
We develop foundation models that reason over complex, linked business data—spanning tables, time series, and graphs. By integrating structural awareness, multimodal inputs, and causal reasoning, our models enable advanced Business AI for analysis, forecasting, and decision-making.
Table representation learning
Learning tabular data representations via table-native and language-based models, integrating business data for advanced reasoning.
Graph neural networks
Using Graph Neural Networks to model relational tabular data, enabling accurate predictions and deeper insights in enterprise AI.
Business knowledge graph
Building enterprise knowledge graphs to enable precise, context-aware queries across diverse business data.
Agentic AI
Building self-improving agents for reliable, goal-driven automation in enterprise systems.
Coding LLM (ABAP)
Empowering enterprise software development with domain-specific ABAP foundation models for intelligent coding assistance.
Track B: Trustworthy AI
Our research develops AI systems that are robust, fair, transparent, and aligned with human values—essential for real-world enterprise use. We focus on robustness, explainability, fairness, privacy, and alignment with domain-specific constraints to ensure reliable and responsible AI deployment.
Differential privacy
We develop efficient deep learning models that save resources and protect privacy.
Data confidentiality
We ensure data confidentiality by protecting structured data and validating privacy through audits and attacks.
Model protection
Analyzing sentiments in text using neural embedding and attention.
Security testing
Enhancing model transparency by making predictions explainable.
Human-Alignment
Extracting data from documents using NLP and computer vision.
Careers
Join us and build the future of Business AI
Work with rich datasets to find machine learning-based solutions to real-world problems in close collaboration with our global network of research partners.
PhD Internship Program
As a PhD Intern you will work with a team of experienced researchers and applied scientists taking on challenges informed by scaling Al methods across and beyond the broad portfolio of SAP's business software. You will have the chance to work with some of the richest data sets available in the world addressing problems that have impact on our customers.
Open Research Positions
(Senior) Data Engineer (f/m/d): SAP Data on RDF Knowledge Graphs (DE)
(Senior) Researcher (f/m/d) - Knowledge Graphs and Generative AI (DE)
Lead Research Engineer (f/m/d) - Foundation and World Models on Structured Data (DE)
Principal Researcher (f/m/d) - Foundation and World Models on Structured Business Data (DE)
(Senior) Researcher (f/m/d) - Foundation and World Models on Structured Business Data (DE)
Publications
Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data
Rishabh Ranjan, Valter Hudovernik, Mark Znidar, Charilaos Kanatsoulis, Roshan Upendra, Mahmoud Mohammadi, Joe Meyer, Tom Palczewski, Carlos Guestrin, Jure Leskovec, ICLR, 2026
Talk, Evaluate, Diagnose: User-aware Agent Evaluation with Automated Error Analysis
Penny Chong, Harshavardhan Abichandani, Jiyuan SHEN, Atin Ghosh, Min Pyae Moe, Yifan Mai, Daniel Dahlmeier, ICLR, 2026
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
Vignesh Kothapalli, Rishabh Ranjan, Valter Hudovernik, Vijay Prakash Dwivedi, Johannes Hoffart, Carlos Guestrin, Jure Leskovec, ICML, 2026
Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky
Ashutosh Hathidara, Julien Yu, Sebastian Schreiber, ACL Findings, 2026
SPRINT: Scalable Secure & Differentially Private Inference for Transformers
Francesco Capano, Jonas Böhler, Benjamin Weggenmann, PETS, 2026
Weishi Wang, Hengchang Hu, and Daniel Dahlmeier, EACL, 2026
SoK: Enhancing Cryptographic Collaborative Learning with Differential Privacy
Francesco Capano, Jonas Böhler, Benjamin Weggenmann, IEEE SatML, 2026
RecPFN: Prior-fitted Networks for In-Context-based Recommendations.
En Zhi Tan, Jia Xiang Lim, Bryan Chew, Tze Minh Ng, Benjamin Yap, Long Papers Track, SIGIR, 2026
Jiyuan Shen, Peiyue Yuan, Atin Ghosh, Yifan Mai, Daniel Dahlmeier, Industry Track, EACL, 2026
Joseph Meyer, Afreen Shaikh, Mohammadi Reza, Dinesh Katupputhur Ramprasath, Karan Paresh, Roshan Reddy Upendra, Tom Palczewski, Mark Li, Summer Symposium, AAAI, 2026
Tao Bai, Zhaochen Li, Hongxin Shao, Daniel Dahlmeier, Industry Track, ACL, 2026
Multi-Level Data Validation for Knowledge Graph Construction Pipelines
Lars Heling, Isaiah Onando Mulang', Hassan el Hajj, Felix Sasaki, Industry Track, ESWC, 2026
MirrorBench: An Extensible Framework to Evaluate User-Proxy
Ashutosh Hathidara, Julien Yu, Vaishali, Sebastian Schreiber, Anil Babu Ankisettipalli, Dataset & Benchmark Track, KDD, 2026
Unified Evaluation of Table Embedding Methods Across Multiple Benchmark Scenarios
Ali Younes, Saeed Ghoorchian, Maximilian Schambach, Johannes Höhne, DATA-FM workshop, ICLR, 2026
Tabular Foundation Model for Generative Modelling
Xiangjian Jiang, Mingxuan Liu, Nikola Simidjievski, Tassilo Klein, Mateja Jamnik, FMSD workshop, ICML, 2026
TableFactory: Generating Semantically Linked tabular Data via Multi-Agent Behavioral Simulation
Mingxuan Liu, Xiangjian Jiang, Johannes Hoffart, Tassilo Klein, FMSD workshop, ICML, 2026
Tassilo Klein, Johannes Hoffart, FMSD workshop, ICML, 2026
Dinesh Katupputhur Ramprasath, Tom Palczewski, Joe Meyer, Roshan Reddy Upendra, Minghua Li, FMSD workshop, ICML, 2026
PLUREL to RDB-PFN: Schema-Guided Synthetic Relational Pretraining
Mohammad Sadeq Abolhasani, Viswanath Ganapathy, FMSD workshop, ICML, 2026
Large-Scale Pretraining unlocks Few-Shot Prediction for Relational Data
Rishabh Ranjan, Vignesh Kothapalli, Harshvardhan Agarwal, Charilaos I. Kanatsoulis, Roshan Reddy Upendra, Tom Palczewski, Carlos Guestrin, Jure Leskovec, FMSD workshop, ICML, 2026
FlexTab: Towards a Flexible Encoder-Decoder Architecture for Tabular In-Context Learning
Marek Polewczyk, Maximilian Schambach, Marco Spinaci, Sam Thelin, Johannes Höhne, FMSD workshop, ICML, 2026
Exploring Differences Between Tabular Enterprise Data and Public Benchmarks
Myung Jun Kim, Maximilian Schambach, Frank Essenberger, André Sres, Johannes Höhne, FMSD workshop, ICML, 2026
Enhancing Tabular Learners with Context-Aware Semantic Embeddings
Günther Schindler, Maximilian Schambach, Johannes Höhne, FMSD workshop, ICML, 2026
Benchmarking Attention for Tabular Foundation Models
Maximilian Schambach, Clemens Biehl, Sam Thelin, FMSD workshop, ICML, 2026
Probing Memorization of Tabular In-Context Learning
Francesco Capano, Jonas Böhler, FMSD workshop, ICML, 2026
Large-Scale Pretraining unlocks Few-Shot Prediction for Relational Data
Rishabh Ranjan, Vignesh Kothapalli, Harshvardhan Agarwal, Charilaos I. Kanatsoulis, Roshan Reddy Upendra, Tom Palczewski, Carlos Guestrin, Jure Leskovec, GFM workshop, ICML, 2026
Beyond Accuracy on RelBench: Item Response Theory Analysis of Relational Deep Learning Benchmarks
Dinesh Katupputhur Ramprasath, Tom Palczewski, Joe Meyer, Roshan Reddy Upendra, Minghua Li, GFM workshop, ICML, 2026
DIPA: Difficulty-Informed Probabilistic Allocation of Test-Time Compute via Training-Free Proxies
Wenyang Hu, Yao Shu, See-Kiong Ng, Bryan Kian Hsiang Low, AdaptFM workshop, ICML, 2026
Incentivizing Black-Box Model Sharing with Fair Rewards and Payoffs
Wenyang Hu, Xinyi Xu, See-Kiong Ng, Bryan Kian Hsiang Low, AAMAS (Extended Abstract), 2026
Isaiah Onando Mulang', Johannes Thaller, Tushar Trivedi, Lars Heling, Felix Sasaki, GENAIK NORA workshop, IJCAI-ECAI, 2026
Parameter-Efficient Vocabulary Expansion via Cross-Lingual Embedding Alignment
Andrew Ivan Soegeng, Muhammad Reza Qorib, Weishi Wang, Daniel Dahlmeier, Hwee Tou Ng, MeLLM workshop, ACL, 2026
Cross-Lingual Consensus: Aligning Multilingual Cultural Knowledge via Multilingual Self-Consistency
Andrew Ivan Soegeng, Patrick Sutanto, Nguyen Tan Sang, MeLLM workshop, ACL, 2026
