A free AI Dev Tools Zoomcamp 2026 starts August 31. Learn using AI developer tools without losing engineering discipline. Register
Season 25, Episode 1

Engineering AI-Powered Data Products | Radovan Bacovic

Show Notes

Timestamps

Click any timestamp to jump to that moment in the video

Transcript

The transcripts are edited for clarity, sometimes with AI. If you notice any incorrect information, let us know.

Data Engineering Career Trajectory and Industry Experience

Alexey: Let's skip the usual intro because we were supposed to start two minutes ago. Do not forget to subscribe to our YouTube channel and check the link in the descriptions for our Slack. There is also a pinned LinkedIn live chat. Use this link for asking questions. I will be monitoring these questions during the interview. (0:00)

Alexey: I want to quickly open this thing. Just one second. I want to make sure that I am actually monitoring these questions. Yes, I have them. Please use this link. Was it the right link? We already have some questions. (0:20)

Alexey: Bako, I guess that might be your nickname. Somebody wrote, "big salute to my friend Bako." (0:47)

Radovan: I have many friends, so I do not recognize who the guy is, but please ping me afterward. We have to talk more. (1:01)

Alexey: Let's start. This week we will talk about engineering AI powered data products and we have a special guest, Radovan. Radovan is a staff data engineer who lives and codes in Novi Sad, Serbia. Did I pronounce the name correctly? (1:07)

Radovan: Yes, Novi Sad is fine. (1:20)

Alexey: I can read Serbian when it is in Cyrillic. When it is Latin, it is a bit more difficult for me. But if I open Serbian Wikipedia, for me as a Russian, it is easy to understand. (1:26)

Radovan: It sounds similar. (1:37)

Alexey: Radovan is also a part of the Data Giants ambassador program and a Snowflake Squad member. I do not know what a Snowflake Squad member is. Maybe you will tell us more about this. You are also a dbt spotlight member. The list is long. (1:37)

Alexey: One interesting thing is that he has been doing data for more than 20 years and he is also a passionate Brazilian jiu-jitsu practitioner. When somebody is doing a martial art, I have the utmost respect for them because I know how much discipline it requires. Do you have a belt? (1:56)

Radovan: Yes, I am a brown belt and aspiring black belt. Next year that will happen and I am so proud of that. It is easier to finish university than get a black belt in jiu-jitsu. There are no shortcuts. (2:13)

Radovan: It reminds me of life challenges all the time. You should be persistent and disciplined. Currently I am not speaking from Serbia, I am speaking from Stockholm, Sweden. I came here for a three day jiu-jitsu camp. It just finished yesterday, so my body is very sore now. (2:25)

Radovan: It was super great and interesting. I can compare it with the IT industry. We have a bunch of lunatics and very nice, funny people eager to help each other. It is the same as our IT industry or my primary life. To get back to your question about Snowflake and dbt, I am part of the ambassador program. (2:42)

Radovan: They have a program for significant members of the community. We share knowledge, go to podcasts, write articles, and solve problems. We help the community to get Snowflake or dbt, the same as any big company. I was trapped in that role for five to seven years in Snowflake. (3:07)

Radovan: They recognized me. In the beginning of June in San Francisco, I took to the stage and gave a very nice talk about AI data products. It was a super great experience. I like to do that. It is my third life to go to the stages and try to explain and give something back to the community. (3:23)

Alexey: It is interesting that you say the word "trapped". This is also what you have in your biography. The line is "he was trapped in the data world for 20 years." Why trapped? Did you accidentally end up in the data world and could not escape it, or what is the story? (3:46)

Radovan: Actually, trapped is in a good way. I do not want to escape. I feel very good here. We will talk about that later, why I am still an individual contributor and find my place in the data world very good and comfortable. Why trapped? Everything happened by accident or coincidence. (4:08)

Radovan: I like what I have been doing for more than two decades now. Somehow I started my career in software engineering but with a main focus on databases. Even at university, it was the same story. I felt in the trap of liking databases because it can be super interesting. (4:25)

Radovan: In that period of time, it was a huge war in the early decades between Java and .NET technologies. It was a total war, and I chose a database role to reconcile both parts and be the guy in the middle contributing there. Things rolled very fast, and we moved from traditional RDBMS systems to big data databases and the AI world these days. (4:38)

Radovan: I saw a lot of things and am more than happy to be here. It is a huge honor to speak on this podcast and try to tell my journey. Maybe someone can be inspired and get shortcuts to avoid the trap I was in. Trapped is in a good way, and I feel comfortable here and want to stay here. (5:02)

Alexey: I do not think it was just Java versus .NET. It was also Java plus an ecosystem. When you use Java, you probably use something like Oracle. With .NET, you probably would use something like Microsoft SQL Server. I am that old, and Java was around even before Oracle acquired Sun. So it was independent. (5:21)

Alexey: Typically, when I think about Java, even before the acquisition, this would be the database of choice. It is for banks and stuff. (5:43)

Radovan: Absolutely. You have a different ecosystem. Are you leaning more on the open source side? You should choose MySQL along with Java. Or you go to Microsoft and use .NET or maybe some other exotic things emerging at that time with web development and PHP. (5:55)

Radovan: Somehow I found myself very comfortably, nice, and most useful on the database side. I also did a lot of application stuff. When data became cool, I started to jump fully into that part. As I said, it was by accident, but it is still very nice to be here. (6:13)

Alexey: What do you mean by jumping into data or being trapped in data? What kind of data things did you do? (6:33)

Full Stack Data Engineer Role Evolution and Responsibilities

Alexey: Was it from the beginning database administration, pipeline creation, detail development, or both? (6:37)

Radovan: This is a good question. What was super interesting is that all of these current things emerged significantly. Actually, I am now a full stack data guy. Before it was like a software engineer full stack. On the data or database side, I always tried to avoid DBA stuff. (6:44)

Radovan: Not because I hate it, but for me it was difficult and not my cup of tea. I think the real value for the business is that you want to produce some value and create some products. Of course, someone needs to do the heavy lifting with the infrastructure. Back then, it was mostly using the data and finding its usage. (7:03)

Radovan: It was not divided per role like ML, data scientist, data engineer, analytics engineer, DBA, infrastructure engineer, or DevOps. It was just one guy doing everything from end to end. Most of the time I was on my own, especially in smaller projects and teams, to do everything. (7:27)

Radovan: The funny story about how I ended up in the data world is that I was a consultant working in a big Dutch telecom company. There was one guy who was retired, a super old guy, and I inherited his role. He told me I was the guy in charge who would do everything. (7:45)

Radovan: I was on my own to do that kind of stuff from end to end administration. It was an Oracle database, some Java, some bash scripts, pipelining, producing reports, and speaking with the business. It was a one-man army in that period of time. How everything emerged in the future is more focused on strategically seeing and overseeing what comes next and trying to organize everything inside the company. (8:05)

Radovan: Currently, I am a principal engineer or data platform lead. I am thinking about what the best way is to organize and create a data platform and how it can work in the future. No one knows what will happen five or ten years from now; you just can guess. You also need to create great building blocks for the future, especially regarding data products, AI, and leveraging data usage. (8:29)

Radovan: I started in one world and went through a few more. Now it is a completely different landscape, not only for me but for all of us. I think no one has the right answer on how this will end up. I have spoken with many people about the feeling with this AI stuff. We are kind of in a fog and keep running all the time, because if you stop, you are done. (9:00)

Radovan: You need to move forward, try and err, fail, move on, and not give up. That is how I see it. (9:27)

Alexey: You have been in the industry long enough to see how much more data we have right now. Later technology was not there, but now everyone uses smartphones. Especially with AI, we are generating so much data. In your role as a principal, you say you need to think about how to organize data platforms. (9:33)

Alexey: How does the scale of data we have to deal with now affect the technologies we need to choose and the data platforms to enable other people to use this data effectively? (9:56)

Radovan: Let's go 10 years back to start answering this question. I was working here in Stockholm in a gambling company and we had some consultants talk to us. They told us about big data emerging and how we can use and leverage it. They told us, based on some research, that data year on year is growing 50%. (10:10)

Radovan: Now it is probably even more. The number of experts is growing only 5% per year. Last year it was one terabyte, now it is 1.5 terabytes. On the other side, the experts who can support the journey are growing only 5%. (10:36)

Alexey: So there are only 5% more experts in the area? (10:55)

Radovan: Yes. I think that part has not changed. The part that has changed is that we have more and more data in exponential growth. The second part of your question is that I have been thinking for the last 10 years how to deal with this and what the problem is. Ten years back, it was a lot of only structured data, like accounting system ledgers and stuff. (11:00)

Radovan: Then other things emerged and now you are actually dealing with semi-structured data, unstructured data, videos, AI data, synthetic data. The million dollar question is how you should build the platform and which route you should take. Whatever you take is probably wrong. The right answer is what is the least wrong path you should take. (11:18)

Radovan: There is actually no silver bullet solution. No one knows what will happen five years from now. In my opinion, you should have ground rules and solid building blocks that can be changed. These days, when I talk to people at conferences and in the community, AI is here and it is not a commodity anymore because anyone can use it. (11:37)

Radovan: What can be an advantage is your data or your company data. That is an asset because everything else is just a commodity. I can open any tool of my choice, type something, and get an answer that can be wrong or right. The next big thing for me is talking with your data and leveraging its usage, especially combining it with AI. (12:02)

Radovan: There is no short or simple answer. Anyone who tries to build a data platform should focus on building a solid foundation and good building blocks. You also need to understand how to bring value to the business. The business will not ask you about the technology you are using, like Java, Python, C, or Scala. (12:28)

Radovan: They will ask for their dashboard and if it is correct. Without that, it does not matter. The two major points for the outcome of your data platform are that people should use your data and your data should be correct. Without that, you will end up in a graveyard with a lot of nice promising applications and data platforms, but nothing will happen. (12:46)

Radovan: That is how you should start from the end. Provide accurate data. (13:07)

Alexey: It was fast but I managed to take some notes. You said nobody knows how to build a data platform. The answer you come up with is probably wrong, but you still need to start somewhere. The important parts are a solid foundation for this data platform and good building blocks. (13:12)

Alexey: You will figure out the answer as you build the platform and more use cases appear. You need to make these building blocks replaceable and flexible. (13:30)

Radovan: They should be flexible because you do not know what will happen in the future and how your business will prioritize. Most companies start building a product, and your data platform should follow the product needs. For example, Slack started out as something completely different than a communication tool. (13:48)

Radovan: You can imagine they pivoted the product and the metrics they tracked. I spoke with one guy and I think it was an internal chat for another product they built. They said, "let's build our chat system for that," and boom, it became Slack as we know it today. (14:06)

Alexey: I see. Back then there were not many good options to use for chat. (14:22)

Alexey: Anyway, I agree with a solid foundation, good flexible building blocks, and the business perspective. (14:37)

Scalable Data Platform Architecture and Business Alignment

Radovan: Yes, I absolutely agree with you and just want to add regarding the business perspective. You should include your business from day one. Many people, including myself, made the mistake of trying to ignore the business. We felt we did not need them and could do something technically. (14:45)

Radovan: You should understand the business first and align with the business needs. (15:05)

Alexey: You need that because you are building the platform for them to use. (15:10)

Radovan: Absolutely. It is not for myself even if I like it. Even if big companies like the red one, the blue one, or Microsoft promise you will have everything in one shot, it will not happen. Sometimes you need to sort out the problem despite the technology. So align with the business needs. (15:10)

Alexey: The red one is Databricks and the blue one is Snowflake. (15:33)

Radovan: Yes. I try to avoid any kind of advertising, which I hate. But you got it. (15:39)

Alexey: This is the first time I heard it. It was like Morpheus with the two pills. (15:46)

Radovan: One funny story just to continue smiling here. You mentioned Oracle, and I was working a lot with Oracle, which is like the red one. AWS created a database and tried to be an MPP database called Redshift. Do you know why it is Redshift? They wanted to shift from Oracle, and Oracle means red, so it is Redshift. (15:53)

Radovan: That is the etymology of the name. They wanted to create a border to show they were not Oracle and were doing something differently. It was very bad, especially regarding managing the cluster, which was a pain in the neck. Sometimes they call it red. I do not know if it is justified, but it is just a funny story. (16:13)

Alexey: Okay, let's get back to building data platforms. A solid foundation, good flexible building blocks, and including the business from day one. What is a solid foundation? Let's say I want to build a data platform and we can come up with a use case. (16:31)

Alexey: It does not necessarily have to be the right choice from day one, but I just want to understand what a solid foundation means. (16:53)

Radovan: Use a typical story like a rising startup that explodes. The business is doing well and earning money. Let's find a startup. (17:06)

Alexey: Let's say they do education. In DataTalks.Club, we do a lot of educational stuff. We teach people and have open courses. A lot of people are enrolling in our courses, so we have data. We are scaling and growing, with more users joining the platform. (17:19)

Alexey: More users are interested in our courses, and we are expanding our course catalog. Everything is good, but actually we do not have any data platform. In my case, all I have is a Django app. I do not have any analytics. I want to start building a data platform. What do I do? (17:39)

Radovan: It is the same problem many companies or communities at your scale have. You generate data and you are scaling, which is a great path and success. But what comes next? You somehow need to do something with this heap of data. You should ask yourself from a business perspective what you want to do with this data. (18:03)

Radovan: That is question number one. What do I want to see and what is the result? The first question is who will use this data? Is it a machine, like an agent, or a visual product like Looker, PowerBI, or Qlik? Let's start very smoothly. (18:29)

Radovan: You need to find one use case. Let's start with education core business data, like your catalog list and product usage. You usually start with product usage. What does usage mean for an educational platform? You might sign up for a few courses and enroll in a couple of them. Or you have a learning path like wanting to be a Python engineer or learn SQL. (18:46)

Radovan: You want to combine those two sources. You have two sources now and one use case called product usage. What the business needs to do is define the metrics to track. (19:11)

Alexey: Let's say completion rate. For each course, we want to optimize the completion rate because we have free open courses. Twenty thousand people can join, but only four hundred will graduate. The funnel is very steep and a lot of people drop out. How can we make sure more people finish? (19:23)

Alexey: This could be a problem. (19:48)

Radovan: This is great because you translated your problem into something quantified in your metrics. I want to know how many people start the course and how many people finish the course. When I divide those two numbers, I get the metric. Let's keep it to one metric; it is super simple but it should scale in the future. (19:54)

Radovan: You have two datasets theoretically speaking: usage and the course list. Usage means a random click here, an enrollment, listening to twenty percent, or finishing the course. How should you use this data? Is it a streaming problem, a batch problem, or do you stick with lambda architecture? People wrongly assume most things should be real time and streaming, which is a pain in the neck. (20:11)

Radovan: For most businesses even today, batch processing is fine because you can survive a one day lag in your data. Your business needs metrics like completion rate and to see how people behave. Then you should discover why everything fails. Now you can drop down to the technical level. This is a very brief simulation. (20:43)

Radovan: You need to store this data somehow. This opens many potential technical problems. Let's say you push or save your data in Postgres or MySQL. (21:08)

Alexey: Nothing else, just Postgres. (21:21)

Radovan: Postgres is a great choice. You can ask yourself if you have so much data or if you can scale. Everything is in Postgres now, and you have plug-in based stuff that was not possible ten years ago. You can use the pg_analytics plug-in. You will hit a wall at some point, but if you want to move fast, you can do everything in Postgres since it is small enough. (21:28)

DuckDB Integration and Premature Database Optimization Avoidance

Radovan: When you see that it starts to slow down when running analytical queries, you need to think about separating them. A major problem people easily overlook is optimizing too early. You can go and buy the red, blue, or Microsoft solution, but actually you are too small and can fit inside Postgres. (21:59)

Radovan: My weapon of choice these days is something called DuckDB. DuckDB is a great Python library and I suggest it to everyone. It is battle tested. In DuckDB, you can easily hook up to your Postgres database and combine it with an S3 bucket or any object storage and do SQL super easy. (22:17)

Radovan: You can host it wherever you want. Let's say I want to be modern and think how I can scale, but it will not scale now. I can use DuckDB to hook up to your table and be ready to go. (22:36)

Alexey: What do you mean by hooking to the table? Do I need some sort of stream of row events that I save to S3, or can DuckDB go directly to my Postgres? (22:47)

Radovan: You can do both, which is the beauty of it. Let's say you have your data inside Postgres and it is just inserted in. You can say DuckDB, I want to attach the Postgres plugin and hook up to my source via URL using a password and database name. You have everything you need inside your DuckDB, but it is raw data. (23:00)

Radovan: You can also combine some third party data from S3 and DuckDB is your junction point. You still build something and have a very fast pipeline because you consume data directly from Postgres without needing to move it from A to B. This simplifies tool chain complexity and is very simple for the beginning. (23:18)

Radovan: The next step is to model the data. We agreed this is a batch problem, and my weapon of choice can be dbt. Not because I am a dbt guy, but dbt in ETL or ELT processing sorts out the 'T', which stands for transform. (23:37)

Radovan: Why is it important to use something like dbt? You do not have many options on the market. dbt is the lingua franca for transformation and it is built using SQL, Jinja, and Python. It is a single source of truth for all transformations. In the past, you started with SQL, created some tables, and a colleague might come and create something else, causing a lot of confusion and redundancy. (23:55)

Radovan: With dbt, you have all the models, definitions, and documentation. You press a button and build your HTML documentation completely interactively outside of the box. Here is the junction point between techniques and business, because the business can help you model the data. Best practice says to organize it as raw, prep, and prod. (24:26)

Radovan: The raw data comes in. Then prep will clean up the data, avoid false positives, and avoid outliers. If you buy Databricks, they will tell you the same story but call it bronze, silver, and gold. The point is the same. Gold is the prod data. (24:55)

Radovan: The business or any other service will tackle only the prod data, which is ready to consume. When you decouple the system like this, you have a great advantage. You can scale easily, you can grow easily, you can put more people on it, and you have a collaborative tool inside one project. Most importantly, you can change underlying data and business logic without touching the prod layer because you model the data all the time. (25:25)

Radovan: You can also introduce more pipelines. Tomorrow you might use web analytics like Google Analytics to see how people behave on your website. This is semistructured data. You can use a RESTful API to go to Google Analytics, drop this data in S3, and just tell DuckDB to connect to this data in another bucket. (25:50)

Radovan: Now you have a single source of truth. If you are very small, you have a zero cost data warehouse because dbt is also open source. If you want to dream big immediately, you will spend hundreds of thousands of dollars for nothing you need. After this meeting, I will share a workshop I attended which covers the same story. (26:19)

Radovan: You will get the full repo with working code and everything we talked about can be tried immediately. Talk is cheap, I just want to give you the full context. I will give the audience free access to this open source course. I think it is the best way to learn, so you can digest what we are talking about today. (26:45)

Radovan: Let's think and do a mental model on how to build everything. We are at the transformation level. The business tells you they want to see a specific metric calculated in a certain way. They want to have data grouped per day, month, year, or quarter. You then run the dbt process and stick with the best practices. (27:09)

Radovan: Best practice means doing everything incrementally. You can do a full refresh all the time if you are super small, but tomorrow you will hit a wall. You can fit everything inside one virtual machine. You can run my workshop locally. When you do the ELT process and transform your data, you have data ready for the business. (27:42)

Radovan: You can run this data once per day or once per hour. The next run will just feed incremental data and you will have a prod layer. With the prod layer, you can go in a couple of directions. Before it was a BI tool where business could click, filter, order, and group the data. Now you can directly create an MCP server. (28:03)

Radovan: You can speak with the agent or monetize this data tomorrow. You can create RESTful API endpoints and charge for them. You can speak with your machine, or maybe an agent can hook up and finish your course. Traditional BI tools are still good, especially for finance when you want to see structured data. (28:29)

Radovan: In my workshop, one interesting use case is using a vector database to speak with my data in plain English. If you have a user or business that does not know SQL, they can speak their native language and it will be translated to SQL. It is all about the text-to-SQL problem. I created the workshop to see the potential of how to build AI based data products. (28:59)

Radovan: You have vectors and AI to translate your text to SQL. But AI should understand your context and definitions. dbt provides documentation in JSON and HTML formats, so you can feed your AI and give it the full context. (29:27)

Alexey: I want to summarize everything we mentioned to make sure I actually understood. Let's say I have Postgres data, and then I have some raw events coming from the system I generate. I save them in S3. Then I have DuckDB plus dbt to process these raw events. (29:56)

Data Pipeline Transformation and Data Modeling with dbt

Alexey: Then I have this prep data or silver layer. Then I can do something with this data to put it on the dashboard. (30:04)

Radovan: Just one correction here. dbt will also provide the gold layer or production layer. All the transformations from raw to prep or bronze to silver, and from prep to prod or silver to gold, we do with dbt or any other transformation tool. (30:27)

Alexey: Eventually, what we have is the data that is ready to use, so we can put it in a dashboard. In our case, we also need to correlate this data with Google Analytics. We can have another pipeline which would pull data from analytics, put it to the raw layer, and we process it and match it. Then we can display both on one single dashboard. (30:46)

Alexey: For the dashboard, we can use any tool. Let's say I can even implement it in Streamlit, or I can take whatever I want. (31:14)

Radovan: Streamlit is a great choice because it is free and open source. You can hook up your AI tool or create an MCP, whatever you want. We are not talking about execution, but about the business process and what you can do with this data. This is a nice recap because you got the point completely. (31:26)

Alexey: What is a data platform in this case? I described a pipeline, but a data platform is something that enables me to execute this pipeline. This is like all the S3 buckets, the things that pull the data from Postgres and Google Analytics, put them to S3, and let me execute DuckDB with dbt on top of that. Would that be the data platform in this case? (31:45)

Radovan: Yes, a very small, very humble, very cheap one. You satisfied the building blocks and basic principles of the data platform. You can think about the ecosystem because you have a multiple heterogeneous environment with many components and technologies inside it. If you go and check the data landscape these days, you can find thousands of tools and you should just use a few of them. (32:07)

Radovan: If you want to be very practical, you can go directly from Postgres and do everything in Postgres with dbt. Or you can do it your own way in bash commands or whatever you want. With this very small and very cheap setup, you have a data ecosystem and your own data platform. What you should attach and plug in the future is up to you and your needs, and it can change over time. (32:40)

Radovan: Most importantly, you can easily scale, change, and pivot. If you put Google Analytics on top of it, or tomorrow you go to your accountant and get an API from your ledger, you can combine it. You have one more pipeline to get the data in, transform the data, and then consume the data. (33:03)

Radovan: The most beautiful part is combining the data to read between the lines. You can discover why your funnel is so steep and why only 10% finish the courses. Then you can be proactive in the future to categorize your customers. For example, if I am a guy from Eastern Europe newly starting in the data world, you could send me a specific nice email. (33:21)

Radovan: If you are from a different part of the world or stage of your career, I will send you a different type of email. You can treat your customer as a first class citizen all the time because you understand the context. That is the most important thing about data. It is not data because of the data, but data to evidence what happened before and also to help shape your future. (33:44)

Radovan: Twenty years back you had ledger data to see how much money you spent or earned in the last two years. Now it is about how to proactively use your data and leverage it to do business in a proper way. (34:10)

Alexey: Would a correct way of defining a data platform be a platform that allows me to ingest data that I need for my business case, transform this data, combine this data, build pipelines, and then display it on the dashboard? (34:24)

Radovan: There is no one good answer, but yes, this is one way you can create a data platform. I spoke a lot with Confluent, who created the Kafka teams. I worked with them for some time, and they try to do this using only Kafka. It is dogfooding. (34:36)

Alexey: The definition of a data platform is a tool that allows us to do this, whatever technology we use. It could be Kafka, it could be Oracle. (34:56)

Radovan: That is it. You can even buy a tool, satisfy the requirements, and say this tool is your data platform. Some companies can live with that. Some companies do not have enough budget for that. You nicely explained and hit the target of what a data platform can be, but it is just one of the options. (35:08)

Modern Data Platform Ecosystem and Infrastructure Orchestration

Alexey: If we go back to the way to approach building that data platform, we need a solid foundation, good flexible building blocks, and a business perspective. We took the business perspective into account while discussing this. (35:31)

Radovan: Solid foundation means figuring out which building blocks you need to solve your problem. If you are doing fraud detection, this approach will not be the right answer. For fraud detection, you have a couple of seconds to see if it is correct or not. You need ML modeling, fast response, and it should be under two seconds or you are done. (35:53)

Radovan: That is a different use case. But if the business says they need to calculate a metric, you look at your building blocks and say you have the data and pipeline to solve the problem. I do not like to use the word pipeline because for me it is just moving data from A to B. We have something more on top of that, like the full transformation process, data persistence, and data encryption. (36:16)

Radovan: You need to satisfy best practices and industry standards. At the end, we need to show this data to someone to consume. Your answer is good, that is one of the options for how you can build your data platform. But the business should drive your requirements and tell you what problem you want to solve. (36:35)

Radovan: A very important point people easily overlook is playing with tools. You should start with the problem. After two decades in data, my advice to anyone is to start with the problem, not with the tool. There are plenty of shiny tools all around that engineers like to play with, but you need to satisfy the business. (36:59)

Alexey: The building blocks in this case for a data platform would be a component that takes data from different sources. We know that right now we have Postgres, but in the future we might have Google Analytics, and later we may switch from Google Analytics to something else. (37:23)

Radovan: Exactly. You need something to orchestrate the process. Airflow is a really good answer here. When you have twenty or two hundred pipelines, how do you orchestrate and see what is failing, what is going fine, and what you need to backfill? You need an orchestration tool on top of that platform. (37:34)

Radovan: I deliberately skipped mentioning that part because we started small and could survive without it. If you want to manage Airflow, there are a couple of options like self-manage, SaaS, or buying a product. Google can manage it, but it can be complicated. (37:52)

Alexey: You want to move fast and solve the problem at the beginning. I can just have a bunch of Python scripts that kind of stick these things together. Would you recommend going this way, or would you suggest using a proper tool from day one? (38:11)

Radovan: I would say it depends on your needs. I am very biased toward open source technologies. When there is a company behind open source, if you scale you can buy it. A good example is dbt. dbt is an integrated project, and you can use it as a command line like putting it in cron to run once a day, which is cool. (38:23)

Radovan: Tomorrow you can use SaaS and pay per developer per month, and dbt will do the heavy lifting. There are plenty of open source tools and open core models. You can get open source but also have enterprise support. The truth is somewhere in between. (38:44)

Radovan: You should also consider your budget. We are living in the cloud era and the AI era, so you always calculate the cost. It is not for free. Do you want to do the heavy lifting or use proprietary solutions? (39:02)

Alexey: Let's say we have this discussion right now on YouTube. YouTube will create subtitles, and what I can do with these subtitles is get the transcript and give it to my Claude. I can say to Claude to implement this and make no mistakes. Do I need data engineers at all right now? (39:15)

Radovan: Absolutely, yes. The problem is it might work perfectly on your machine based on the code Claude built, even if you gave it all the details, high level explanations, your context, and hooked it up to your Postgres. But how should you scale it, deploy it, maintain it, and most importantly, handle security? (39:38)

Radovan: You need a data engineer or a data guy more than ever. Claude can do the heavy lifting for you, but you need to maintain the deployment, development, testing, and security processes. We did not tackle that yet because we just focused on the high level components and building blocks of the data pipeline. From prototype to reality or production is a long way through. (40:16)

Radovan: From my experience, AI code is super good to give you a prototype to persuade the business with a shiny report. But your way to production in reality is really long, and you need a data guy to bridge that gap. We created a based data product on a production scale, and the prototype was done in two days. Six months later, we struggled with some basic things because it is super difficult. (40:43)

Radovan: You might have sensitive data, and if you are in the EU, it is under GDPR compliance. You must be legally compliant, which is a pain in the neck. If you want to run quickly, put everything into code, but then you need someone to revisit that and automate it. You need DevOps, pipelines, testing for the Python code, compliance, and security checks. (41:16)

Radovan: You need the human in the loop to verify all the stages, correct what AI did not understand properly, and understand your context. I would say an AI tool is blind in that region. (41:42)

AI Powered Data Products and LLM Pipeline Automation

Alexey: How then do we use AI to actually build the data platform? Was my idea of taking the transcript and having a prototype the right direction, or should I approach it differently? (41:58)

Radovan: There are two branches in how I see this stuff these days. One branch is using AI to speed up your development, get the code, and get the building blocks. The second part is actually building AI products based on your data. You can use both branches. You can use AI generously to help you speed up your development, deployment, and security. (42:16)

Radovan: But you should revisit everything and have a human in the loop to check what is missing, what is wrong, and if it is good enough. On the other side, with the data you have when you build a data platform, you should create AI based data products. You can create an MCP server, use LLM stuff to talk with your data, or create an agent to give suggestions to users. (42:45)

Radovan: You could imagine someone going to your website saying they are Radovan, 44 years old, from Serbia, with specific skills, wanting to become an LLM guy. The agent could then identify the gap and offer suggestions. Your approach on how you want to start is perfectly fine, but with the remark that you have a long way through to get to production. If you just do AI code, it will not end up in a good way. (43:09)

Alexey: I need to treat it as a prototype. By no means do I treat it as ready for production. (43:46)

Radovan: Absolutely. You should push it to production with the generous help of AI on each stage, but you need milestones where you want to check things. Did you get the data from Postgres in a good way? Did you put some private data that shouldn't be there? You need to clean up the data and tell AI to create a Python script to clean up the data on the fly. (43:52)

Radovan: Do you need to mask some data, like credit cards, to comply with compliance rules? How can you deploy it, what is a good strategy for deployment, and how to deal with bringing more people on? Many questions are open, and a data guy can help you push things to production from a prototype much quicker than you could on your own. (44:09)

Alexey: I am not a data engineer, so for me it is good to have this conversation to know what kind of milestones and checkpoints there are. I need a data person who would tell me that I can do that, but to think about specific milestones. At each of these milestones, you have to be present and check that the input and output from the system actually match what you expect. (44:38)

Radovan: Absolutely. Let's imagine a scenario where you are a big company and want to buy a large proprietary database, like the blue one, Snowflake. You put everything in place, create the tables, and it is running fine. Then you need to pay one hundred thousand dollars to Snowflake because you overutilized your warehouses. (45:08)

Radovan: You have security, ingestion, and streaming out of the box, and the heavy lifting is on their side, but your cost can explode. You then need to deal with the cost. For that kind of fine-tuning, experience, and battle tested stuff, you still need a data guy. I am currently exploring another MPP database that is open source. (45:31)

Radovan: AI helped me in really two paragraphs to understand its concepts and how it is different from Snowflake. Since I know Snowflake but not that tool, I got a really quick answer in ten minutes just before our conversation. That is how AI can help you. But still, the decision and judgment on how you should define the data strategy and which tools you should use is mine. (45:54)

Radovan: AI can help you compare, explore, and discover, and give you the building blocks to start. (46:22)

Alexey: We have quite a few questions, and I want to take a step back. You mentioned that there are two ways we can use AI: speeding up development and building AI products. The first question we have is about the first category, how we can use AI to build data platforms. You mentioned that one of the checks we need to do when building this is thinking about masking PII data. (46:35)

Alexey: The question we have is how to think about adding guardrails in production for data processing pipelines that have PII or confidential data. How do I approach this? (47:00)

Radovan: In large companies, you have something called data governance. It can be a team, one person, or any person in the room taking care of data governance. Data governance means you should categorize, organize, tag, clean, and define your data. It is a meta layer on top of your data platform. (47:24)

Radovan: We did not tackle that yet because we started building and playing with one use case. Somehow in the future, you need to have an answer to this. If you scale your business 100 times, you need data governance. If you work your business on two or three continents, you need different data governance because what is allowed in the US is not allowed in the EU. (47:46)

Radovan: It comes with multiple problems, and you should slowly start to think from day one about security and data governance. You need to tag the data. This dataset belongs to this source, it takes care of finance data, or this dataset contains PII or MNI data, which is an asset for your company. By default, my advice is do not put any kind of sensitive data inside your data warehouse. (48:04)

Radovan: However, without this data, you cannot generate great insights. For instance, how should I save your email? It is a trivial question but comes with drama because you may not want to expose who the person is inside the data warehouse. Do you really need the email, or do you just need a category or domain of that user? (48:39)

Radovan: Many questions are open. Long story short, think about data governance and start small because you will be busy with other things. But think about compliance, governing the data, and security around your data platform. You should start from day one, but very small, because you are busy bringing value to the business. At one point, you need to consider that very seriously. (49:03)

Enterprise Data Governance and Pipeline Security Best Practices

Alexey: We were talking about security, and somebody mentioned in the live chat that security is the key. What do we mean exactly by security? Is it that data cannot be hacked and data cannot be leaked, or is that just one aspect? (49:29)

Radovan: Security is a pain in the neck for data guys these days. In software engineering, you create an application and have typical security. In the data world, you need to protect your pipeline, your tools, and your Python libraries. On top of that, you need to protect the data, so you have a much bigger problem. (49:42)

Radovan: You could have vulnerable code inside your data ecosystem, and your data can breach. Generally speaking, a data breach will happen to everyone at some point. We just need to protect ourselves and try to delay the day as much as possible. Regarding security, stick with the best practices. (50:11)

Radovan: Do not hardcode your credentials, protect your cloud, and implement the principle of least privilege. Give people only what they want to see, especially when we are talking about data. (50:41)

Alexey: This security is also an aspect of data governance that we talked about, regarding who should have access. (50:54)

Radovan: Absolutely. We are talking about security, like removing vulnerable Python libraries and trusting only verified publishers. On the other side, data governance and security overlap, especially concerning who can see what. I always stick with the least privileges principle. As a data guy, you can probably see everything and are a kind of insider. (51:00)

Radovan: You should treat it differently because it is very sensitive. You could walk away, sell that to a competitor, and be rich. Another person building the platform can see and access everything, but you should give the business access to the production layer only. You should also do fine granulation inside the production layer. (51:26)

Radovan: Asia can see Asia data only and cannot see salaries because of known reasons. You should think about the granulation of privileges on the data level and prevent any kind of breach as much as possible. You should do encryption at rest and in transit, and stick with the best practices. Apply this to your policy and politics, be very strict, and be paranoid. (51:48)

Alexey: Another question is what is the fastest way to learn dbt these days? If I want to learn a technology, I ask an agent to implement it, then I try to learn how it did it, and ask questions. In your experience, is it the fastest way to do this, or are there other ways? Should I take a course instead, or how would you recommend approaching learning dbt specifically, and technologies in general? (52:22)

Radovan: What you said is really a shortcut, kind of reverse engineering in the AI era, and I like that. That comes with a price. If you are an experienced guy like you and me, it is the fastest way. You understand the basics, the building blocks, and the technology, and you have failed many times. (52:54)

Radovan: On the other side, if people are starting in this world, they should learn the basics first. It is like jiu-jitsu. If you try to beat someone who is experienced, you will be beaten. Learn the basics first, then start sparring, then start competing. If you are a black belt in jiu-jitsu, you can spar with anyone because you know how to protect yourself. (53:18)

Radovan: The same principles apply here for dbt, Python, SQL, or whatever you get. For us as experienced guys, we need to move fast and do not have time to spend two weeks learning something. I usually do reverse engineering. I will share a link to my workshop, and a small portion of that is dbt. (53:42)

Radovan: I told Claude I have this dataset, please create something nice sticking with the best practices of raw, prep, and prod. It created it, it was working, and I asked it to create tests for that. Then I reverse engineer to see what it did, fix things, put more tests, remove what I do not need, and with some short intervention, I have everything I need. (54:05)

Radovan: Then you can learn, but I would say do not skip the basics. The trap these days for most people is saying "AI, let's do it," and they skip the basics. They end up with a ridiculous application and do not understand what is going on. If you understand what is going on, you will succeed. AI is an equalizer; if you are smart, you will be smarter, and if you are dumb, you will be dumber. (54:25)

No Code Data Integration Tools and DataOps Methodology

Alexey: What do you think about no-code solutions when it comes to data pipelines and AI engineering in general? Do you think it is a good idea? Should we avoid them or master them? What is your opinion on them? (54:55)

Radovan: You should develop critical thinking on everything, including no-code solutions. It is good because you allow people outside of technology to bring ideas into reality. But on the other side, if I create a mess now, five years from now you will have a huge graveyard of messy applications. (55:14)

Alexey: I already have that. I do not need to wait for five years. (55:39)

Radovan: You got the point now. What to do with this? If you jump into a no-code solution, you have something you should monetize, and then Kubernetes goes down. What is your next step, to give up or to say something to customers? It should be a process way. (55:44)

Radovan: No code is fine, but you should put it under a process. In any industry you have people, tools, and processes. No-code solutions are the tools, but you still have people and should have processes. If you have only people and tools but no processes, it is not promising. (56:01)

Alexey: DevOps advocates that you have to have these processes and automation. (56:24)

Radovan: It is a must. Five years ago I was working at GitLab, and I do DevOps on my own in the data world, something we call DataOps. You should be capable of understanding, implementing, applying, and maintaining that part. We did not touch that because our example was super small, but you also need to automate this. (56:31)

Radovan: You cannot hire five DevOps engineers only to maintain the data platform. You should maintain it on your own. You see the overlapping of a couple of disciplines: AI, data, software engineering, and LLMs. With no code, it will just create a house and nothing else. You should stick with the process, and when I talk about process, it is not only technical process but business process as well. (56:50)

Radovan: When you have an idea and want to put it in production, what is the process? Where should you write it down, how should you deal with the code, how should you sell your idea to leadership, who will budget it, and who will own the project? There are many questions before the coding actually starts. (57:20)

Alexey: I have nice experience with Zapier as a no-code tool for AI. It allowed me to automate many things, but at some point it just became too expensive because of the way they charge. There is this risk of getting something fast, but when you start scaling, you need to either pay more or think how to migrate. (57:34)

Alexey: Then you are back to square one. Now with Claude, it is faster to do something than Zapier two years ago. I can just ask and it will write code that will be cheaper potentially, but potentially also blow up at some point. (58:02)

Radovan: Engineering is about tradeoffs. As I told you in the beginning, there is no right answer; any decision is wrong, let's choose the least wrong decision we can live with. It is about trade-offs, what you should do, who will manage that, and if the heavy lifting is on your side. Do you need to hire people or can it scale? Always think how you can scale your business 100 times without pain. (58:25)

Alexey: There are still a few questions left that we did not cover, but you mentioned the workshop you did, and I think some of the questions will be answered there. Please share the link. I will put this link in the description of this video. Thanks, I really enjoyed this discussion. (58:51)

Alexey: I got a lot of information from you. I took a lot of notes and it was really great talking to you. Thanks a lot for sharing all these things with us today. (59:16)

Radovan: Thanks a lot to you and the audience. I am always very close to the open source community and like to share my knowledge, so this is a good megaphone for me to share my experience. Any questions you have, connect with me on LinkedIn or drop me an email, I am more than happy to help. (59:27)

Radovan: I will share the workshop link with you. We have many more questions and details there. Huge thanks to all of you. Let's keep in touch and help each other go through this uncertain time with AI and see what will happen. Have a great day. (59:46)

Alexey: If we just type in LinkedIn, Radovan Bacovic, we will find you. (1:00:02)

Radovan: I am probably the only person there with this name and surname. Easy to find me. (1:00:09)


DataTalks.Club. Hosted on GitHub Pages. Built with Rustkyll. We use cookies.