curl "https://api.ideardev.com/api/v1/scrapings?page=1&per_page=20" \ -H "Authorization: Bearer $IDEARDEV_TOKEN"
Scraping
A scraping is a saved template: a URL, the steps to reach the data and the fields to extract.
apihttps://api.ideardev.com/api/v1api3https://api3.ideardev.com/api/v1curl "https://api.ideardev.com/api/v1/scrapings/44" \ -H "Authorization: Bearer $IDEARDEV_TOKEN"
curl -X POST "https://api.ideardev.com/api/v1/scrapings" \
-H "Authorization: Bearer $IDEARDEV_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"url": "https://portal.gov.co/consulta",
"scrap_type": "html",
"keywords": "name, status"
}'curl -X PUT "https://api.ideardev.com/api/v1/scrapings/44" \
-H "Authorization: Bearer $IDEARDEV_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"scraping": {
"keywords": "name, status, date"
}
}'curl -X DELETE "https://api.ideardev.com/api/v1/scrapings/44" \ -H "Authorization: Bearer $IDEARDEV_TOKEN"
POST/scraping_run
Runs the scraping and returns the data in the same response. It can take more than 50 s.
api3.ideardev.com · runs
idbodycurl -X POST "https://api3.ideardev.com/api/v1/scraping_run" \
-H "Authorization: Bearer $IDEARDEV_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"id": 101
}'Response
{
"message": "ok",
"scraping": {
"id": 101,
"status": "completed"
},
"data": [
{
"name": "…"
}
]
}POST/scraping_start
Tests the configuration against the site before saving it. The extraction is charged.
api3.ideardev.com · runs
urlbodycurl -X POST "https://api3.ideardev.com/api/v1/scraping_start" \
-H "Authorization: Bearer $IDEARDEV_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"url": "https://portal.gov.co/consulta"
}'The agent downloads the page, proposes the full configuration and tests it.
api3.ideardev.com · runs
urlbodypromptbodycurl -X POST "https://api3.ideardev.com/api/v1/agent_scraping_builder" \
-H "Authorization: Bearer $IDEARDEV_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"url": "https://portal.gov.co/consulta",
"prompt": "name and status of the filing"
}'Validates an edited configuration without saving it. Returns status, issues, fields and preview.
api3.ideardev.com · runs
curl -X POST "https://api3.ideardev.com/api/v1/scraping_validate" \ -H "Authorization: Bearer $IDEARDEV_TOKEN"
Detects the candidate containers of a URL and the engine that served the page.
api3.ideardev.com · runs
urlbodycurl -X POST "https://api3.ideardev.com/api/v1/scraping_containers" \
-H "Authorization: Bearer $IDEARDEV_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"url": "https://portal.gov.co/consulta"
}'Data extracted by the scraping. This is the endpoint to read it from your system.
api.ideardev.com
idpathcurl "https://api.ideardev.com/api/v1/data_scraping/by_scraping_id/44" \ -H "Authorization: Bearer $IDEARDEV_TOKEN"
curl "https://api.ideardev.com/api/v1/scraping_export/44/json" \ -H "Authorization: Bearer $IDEARDEV_TOKEN"
curl "https://api.ideardev.com/api/v1/scrapings/44/availability" \ -H "Authorization: Bearer $IDEARDEV_TOKEN"