# Data Struggling > Just another tech blog... ## Posts - [Surviving the Promptpocalypse: A Guide to Better LLM Conversations](https://datastruggling.com/surviving-the-promptpocalypse-a-guide-to-better-llm-conversations/): Lee Boonstra just released a beautiful 68-page Prompt Engineering whitepaper that I will summarize here. You’ve surely heard the buzz about LLMs and Agentic AI over the last couple of years. Think of them as that brilliant, slightly chaotic intern you once had. Super smart, tons of potential, but if you mumble your instructions, you’ll get a beautifully written report on the wrong topic. Prompt engineering is basically learning to talk to this intern, your instruction manual for the AI. You’re basically giving it a starting point and hoping it predicts the next words in a way that makes sense for *your*... - [A summary of how AI has progressed in the last 5 years and current challenges (by ChatGPT)](https://datastruggling.com/a-summary-of-how-ai-has-progressed-in-the-last-5-years-and-the-current-challenges-we-must-overcome-written-by-chatgpt/): Hey there, fellow data strugglers! It’s no secret that the field of artificial intelligence (AI) has been making massive strides over the last few years. Let’s take a quick trip down memory lane and see how things have progressed since 2017. First up, we’ve seen the release of some pretty impressive AI algorithms that have really pushed the boundaries of what we thought was possible. One such algorithm is GPT-3, a language model developed by OpenAI that’s capable of writing human-like text. We also can’t forget about AlphaGo, which became the first computer program to beat a world champion in the... - [4 reasons for Agile in Analytics](https://datastruggling.com/4-reasons-for-agile-in-analytics/): Working on several projects as a Data Science consultant, I’ve realized the need to spread the word about project planning in this field. This is neither an Agile apology nor an open letter criticizing project managers (PMs) who prefer other methodologies. It’s more of a post trying to help analysts struggling with deadlines when their manager’s Gantt diagram is not really helpful. There are tons of different project management methodologies (see a nice article here) and different ways to apply Agile (doing or being Agile). However, in this article, I’ll focus on a comparative Waterfall vs. Agile (doing) approaches, as, in my experience,... - [Struggling with Hive... What can I do?](https://datastruggling.com/struggling-with-hive/): If you are in a Big Data project, you may have experienced how slow is Hive to JOIN a couple of tables of few TBs (well, even GBs being honest). The first option always appears to be using PARQUET as your default storage engine and then when a query is too heavy, use Impala to process it. Wait, is it that easy? Well, it depends on many factors. First, an updated version of all these tools will help a lot. If it’s not updated, or your company has its own vendor patches (Cloudera, Hortonworks…) for cyber security purposes, then start praying…... - [Working in a Big Data Project using the terminal](https://datastruggling.com/terminal-commands-in-a-big-data-project/): So, you are just landing in a big data project. Everybody knows how to use HDFS except you. All the data is in such a big cluster and you don’t know how to access to it. You are not really into graphic interfaces, so you don’t really enjoy Cloudera Workbench/ HUE / Zeppeling or other of the tools that you are likely to have in your company. Don’t Panic! See here some quick advices and tricks for using your new environment through the terminal.   Create directory in HDFS in terminal: hdfs dfs -mkdir hdfs://path See directory in terminal hdfs dfs... - [Scraping stock prices using Alpha Vantage and Google Finance](https://datastruggling.com/scraping-stock-prices-using-alpha-vantage-and-google-finance/): Stock price scraping can be a nightmare if the APIs you’re trying to use are not up to date. Few months ago I was looking for free sources to obtain one-min-level data. Apart of having troubles with the Yahoo Finance API (apparently non up-to-date by then) and having to tweak some code samples in GitHub to scrape Google Finance, I found the new and shining provider of real-time stock prices, Alpha Vantage. So I decided to develop an script to download data once a week for backtesting. The Alpha Vantage API is as straightforward to use as seen below. It returns a... - [Flattening complex XML structures into Hive tables using Spark DFs](https://datastruggling.com/flattening-xml-structs-into-hive-tables-using-spark/): A couple of months ago in work we faced an issue where we got XML files with nested structs in structs and arrays (with also structs in them). Normally we always face these issues in Hive. Our ETL guy ingests the XML in HDFS in a Yarn cluster in AVRO format. Then we do SQL using Hive no matters what… The thing here is that our Data Engineer basically discovered that Spark would take about 20 minutes roughly on performing an XML parsing that took to Hive more than a day.  Basically she tested the same job in Hive (exploding multiple arrays) and... ## Pages - [Cookie Policy](https://datastruggling.com/cookie-policy/): This page provides comprehensive information about how we use cookies on our website to enhance your browsing experience, improve website performance, and deliver personalized content. Cookies are small text files that are stored on your device when you visit our site. They help us understand how visitors interact with our website, allowing us to offer a smoother and more efficient user experience. In the table below, you will find detailed information about each type of cookie we use, their purpose, and how long they remain on your device. We are committed to respecting your privacy and providing transparency about the data... - [About](https://datastruggling.com/about/): [vc_row css_animation=”none”][vc_column width=”1/2″ css=”.vc_custom_1510335858673{padding-right: 40px !important;}”][vc_single_image image=”1037″ img_size=”full” css_animation=”fadeInLeftBig”][/vc_column][vc_column width=”1/2″ css=”.vc_custom_1510345929871{padding-top: 40px !important;padding-right: 40px !important;padding-left: 40px !important;}”][vc_column_text css_animation=”bounceInRight” css=”.vc_custom_1510336953869{margin-bottom: 0px !important;border-bottom-width: 0px !important;padding-bottom: 0px !important;}”] WHO AM I? [/vc_column_text][vc_column_text css_animation=”bounceInRight” css=”.vc_custom_1747412462762{margin-top: 0px !important;border-top-width: 0px !important;padding-top: 0px !important;}”] Dr. Andrés L Suárez-Cetrulo Data Science Architect and Geek bitten by the Travel bug  I currently lead AI projects applied to industry in Ireland’s Centre for Artificial Intelligence, and I have been working and researching in the Artificial Intelligence and Data Science space since 2012. Besides my research focus on model efficiency, robustness, and trustworthiness, I love travelling, singing, drinking mate 🧉, and lifelong... - [Post Page](https://datastruggling.com/post-page/) - [Sliders](https://datastruggling.com/sliders/): [vc_row full_width=”stretch_row_content” css=”.vc_custom_1492526516622{margin-top: 25px !important;margin-right: 25px !important;margin-bottom: 25px !important;margin-left: 25px !important;}”][vc_column][vc_empty_space height=”25px”][vc_empty_space height=”25px”][vc_empty_space height=”25px”][/vc_column][/vc_row] - [Home](https://datastruggling.com/): [vc_row css=”.vc_custom_1518089562344{margin-top: 50px !important;margin-bottom: 80px !important;}”][vc_column][/vc_column][/vc_row][vc_row css=”.vc_custom_1496413876832{margin-bottom: 80px !important;}”][vc_column][vc_empty_space height=”45px”][/vc_column][/vc_row] - [Sample Page](https://datastruggling.com/sample-page/) ## Optional - [Agent (MCP protocol)](websites-agents.hostinger.com/datastruggling.com/mcp) [comment]: # (Generated by Hostinger Tools Plugin)