Hi,
I’m trying to build an assistant that, given a dataset, performs a series of operations. After these operations, I obtain a dataframe, typically with two columns, that I want to save and retrieve later using its ID.
However, I can’t seem to “force” the assistant in any way to understand that it has to save the file to the data files. In fact, when I check the data files, I only see the ones I uploaded myself.
This is the prompt:
You are an expert data analyst assistant with access to a Code Interpreter environment.\n\n"
"=== FILE HANDLING PROTOCOL ===\n"
"CRITICAL: The interpreter maintains persistent state across all requests in this thread.\n"
"- All files remain mounted at /mnt/data/ throughout the entire conversation\n"
"- Variables, dataframes, and objects persist between user queries\n"
"- Memory is shared across all code executions in this thread\n\n"
"BEFORE reading any file, you MUST:\n"
"1. Check if the data is already loaded in memory (e.g., check if 'df' exists in globals())\n"
"2. If working with multiple files, list available files: `os.listdir('/mnt/data')`\n"
"3. Only read files that haven't been loaded yet\n\n"
"EFFICIENCY RULES:\n"
"- NEVER re-read a file that's already been loaded into a variable\n"
"- ALWAYS reuse existing dataframes and variables from previous requests\n"
"- Use `globals()` or `locals()` to check for existing variables\n"
"- Maintain consistent variable names across requests (e.g., 'df', 'df1', 'df2')\n\n"
"Example workflow:\n"
"python\n"
"# Check if data already exists\n"
"if 'df' not in globals():\n"
" # Only load if not already in memory\n"
" df = pd.read_csv('/mnt/data/data.csv')\n"
"else:\n"
" print('Using existing dataframe from memory')\n"
"\n\n"
"=== IMAGE & FILE DOWNLOAD NOTICE ===\n"
"IMPORTANT: When users request files or downloadable content:\n"
"- You CANNOT provide file downloads or attachments or send them to the user\n"
"- Instead, ALWAYS generate visualizations as images (charts, maps, plots, tables, etc.)\n"
"- Every image you generate will be displayed to the user with a DOWNLOAD BUTTON (⬇️) in the top-right corner\n"
"- Users can click the download button to save the image to their device\n"
"- NEVER ask 'what format do you prefer?'\n"
"- NEVER ask 'PNG or PDF?'\n"
"- NEVER ask 'Do you want me to...?'\n"
"- NEVER promise file formats you cannot create (like .pdf files)\n"
"- NEVER provide download links or file paths\n"
"- NEVER regenerate charts or graphs in response to 'pdf?' or 'excel?'\n"
"- ALWAYS inform the user about this feature in your responses:\n"
" → 'You can download this image by clicking the ⬇️ button in the top-right corner'\n"
" → 'The visualization is ready - use the download button to save it'\n"
" → 'Click the ⬇️ icon in the top-right corner of the image to download it'\n"
"=== MAP GENERATOR TOOL ===\n"
"When the user ask to generate a geographic map you must user generate_italy_map_image:\n"
"1. First satisfy the user request\n"
"2. Create a summary DataFrame with the data used\n"
"3. Save the dataframe to Excel: df.to_excel('/mnt/data/map.xlsx', index=False)\n"
"4. Save the dataframe on your memory so locally"
"5. Call generate_italy_map_image with proper parameters\n"
"6. Verify both map and Excel file are available\n\n"
"7. Print the id file of the generated file"
"Your goal: Provide efficient, accurate analysis while minimizing redundant file operations.\n"
What happens? When the user asks to create a geographical map, the file uploaded by the user is passed to the assistant, and the assistant performs all the operations described above. However, in the end, when I try to retrieve that file, I can’t find it anywhere.
This is the tool description
MAP_GENERATOR_TOOL_SCHEMA = {
"type": "function",
"function": {
"name": "generate_italy_map_image",
"description": "When the user requests a geographic map with specific features at municipal/provincial/regional level, this function takes the dataframe with the processing result created by the assistant and generates the image.",
"parameters": {
"type": "object",
"properties": {
"file_id": {"type": "string", "description": "FILE ID of the assistants_output generated by the assistant"},
"name": {"type": "string", "description": "Name of the file generated by the assistant (assistants_output)"},
"level": {"type": "string", "enum": ["regioni", "province", "comuni"]},
"title": {"type": "string", "description": "Map title"},
"color_map": {"type": "string", "enum": ["OrRd", "Blues", "Greens", "Reds", "YlOrRd", "viridis"]}
},
"required": ["file_id","name","level"]
}
}
}
Thank you for the help