AutomationFlowsWeb Scraping › Firecrawl URL List Mini-batch to Resilient Analyzer

Firecrawl URL List Mini-batch to Resilient Analyzer

04 - Firecrawl URL List Mini-Batch to Resilient Analyzer. Uses googleSheets, httpRequest, executeWorkflowTrigger. Event-driven trigger; 42 nodes.

Event trigger★★★★★ complexity42 nodesGoogle SheetsHTTP RequestExecute Workflow Trigger
Web Scraping Trigger: Event Nodes: 42 Complexity: ★★★★★ Added:

This workflow follows the Execute Workflow Trigger → Google Sheets recipe pattern — see all workflows that pair these two integrations.

The workflow JSON

Copy or download the full n8n JSON below. Paste it into a new n8n workflow, add your credentials, activate. Full import guide →

Download .json
{
  "name": "04 - Firecrawl URL List Mini-Batch to Resilient Analyzer",
  "nodes": [
    {
      "id": "fc000000-0000-0000-0000-0000000000n1",
      "name": "Overview Note RU",
      "type": "n8n-nodes-base.stickyNote",
      "typeVersion": 1,
      "position": [
        -360,
        -160
      ],
      "parameters": {
        "content": "## 04 \u2014 Firecrawl: \u0441\u043f\u0438\u0441\u043e\u043a URL (\u043c\u0438\u043d\u0438-\u0431\u0430\u0442\u0447) \u2192 \u0443\u0441\u0442\u043e\u0439\u0447\u0438\u0432\u044b\u0439 \u0430\u043d\u0430\u043b\u0438\u0437\u0430\u0442\u043e\u0440\n\n\u041e\u0431\u0440\u0430\u0431\u0430\u0442\u044b\u0432\u0430\u0435\u0442 \u0420\u0423\u0427\u041d\u041e\u0419 \u0441\u043f\u0438\u0441\u043e\u043a \u0438\u0437 3\u20135 URL \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u043e\u0432 \u043f\u043e \u043e\u0434\u043d\u043e\u043c\u0443.\n\u041d\u0415 \u0440\u0430\u0441\u043f\u0438\u0441\u0430\u043d\u0438\u0435, \u041d\u0415 crawl, \u041d\u0415 batch-\u044d\u043d\u0434\u043f\u043e\u0438\u043d\u0442, \u041d\u0415 \u0431\u043e\u043b\u044c\u0448\u043e\u0439 \u043f\u0430\u0440\u0441\u0438\u043d\u0433.\n\n\u0427\u0442\u043e \u0434\u0435\u043b\u0430\u0435\u0442 \u043d\u0430 \u043a\u0430\u0436\u0434\u044b\u0439 URL:\n1. \u041d\u043e\u0440\u043c\u0430\u043b\u0438\u0437\u0443\u0435\u0442 URL \u0438 \u043f\u0440\u043e\u0432\u0435\u0440\u044f\u0435\u0442 \u0434\u0443\u0431\u043b\u0438\u043a\u0430\u0442 \u0432 \u043e\u0442\u0434\u0435\u043b\u044c\u043d\u043e\u0439 \u0432\u043a\u043b\u0430\u0434\u043a\u0435 url_registry (\u043f\u043e normalized_source_url) \u0414\u041e \u043b\u044e\u0431\u044b\u0445 \u0442\u0440\u0430\u0442.\n2. \u0414\u0443\u0431\u043b\u0438\u043a\u0430\u0442 (\u0438 force_reprocess=false) \u2192 skipped_log, parse_method=dedup_source_url, \u0411\u0415\u0417 Firecrawl/Claude (0 \u0441\u0442\u043e\u0438\u043c\u043e\u0441\u0442\u0438).\n3. \u041d\u043e\u0432\u044b\u0439 URL \u2192 Firecrawl (POST /v2/scrape, markdown) \u2192 \u043e\u0447\u0438\u0441\u0442\u043a\u0430 (\u0431\u0435\u0437 \u043a\u0430\u0440\u0442\u0438\u043d\u043e\u043a/svg, \u0440\u0435\u043b\u0435\u0432\u0430\u043d\u0442\u043d\u043e\u0435 \u043f\u0435\u0440\u0432\u044b\u043c) \u2192 \u0437\u0430\u043f\u0438\u0441\u044c-\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a (text_context \u2264 3500).\n4. \u2192 \u0442\u043e\u0442 \u0436\u0435 \u0443\u0441\u0442\u043e\u0439\u0447\u0438\u0432\u044b\u0439 \u0430\u043d\u0430\u043b\u0438\u0437\u0430\u0442\u043e\u0440 (Claude \u2192 \u0440\u0430\u0437\u0431\u043e\u0440 \u2192 \u0440\u0435\u043c\u043e\u043d\u0442 \u043f\u0440\u0438 \u0441\u0431\u043e\u0435 \u2192 \u0434\u0435\u0442\u0435\u0440\u043c\u0438\u043d\u0438\u0440\u043e\u0432\u0430\u043d\u043d\u044b\u0439 fallback \u043f\u043e \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u0443 \u2192 \u043d\u043e\u0440\u043c\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044f \u2192 \u043c\u0430\u0440\u0448\u0440\u0443\u0442).\n5. \u2192 \u043d\u0443\u0436\u043d\u0430\u044f \u0432\u043a\u043b\u0430\u0434\u043a\u0430 Google Sheets (Sheet Name = route).\n6. \u041f\u043e\u0441\u043b\u0435 \u043e\u0431\u0440\u0430\u0431\u043e\u0442\u043a\u0438 (\u043d\u0435 \u0434\u0443\u0431\u043b\u0438\u043a\u0430\u0442) \u2192 \u0441\u0442\u0440\u043e\u043a\u0430 \u0432 url_registry (10 \u043a\u043e\u043b\u043e\u043d\u043e\u043a).\n\nFirecrawl \u0443\u043f\u0430\u043b/\u043f\u0443\u0441\u0442\u043e \u2192 technical_errors \u0411\u0415\u0417 Claude (\u043d\u043e \u0432\u0441\u0451 \u0440\u0430\u0432\u043d\u043e \u043f\u0438\u0448\u0435\u0442\u0441\u044f \u0432 url_registry). \u0421\u0431\u043e\u0439 \u043e\u0434\u043d\u043e\u0433\u043e URL \u043d\u0435 \u043e\u0441\u0442\u0430\u043d\u0430\u0432\u043b\u0438\u0432\u0430\u0435\u0442 \u043e\u0441\u0442\u0430\u043b\u044c\u043d\u044b\u0435.\n\u041f\u0440\u0430\u0439\u043c\u0435\u0440\u0438+\u0440\u0435\u043c\u043e\u043d\u0442 \u043d\u0435 \u0434\u0430\u043b\u0438 JSON, \u043d\u043e \u22655 \u0441\u0438\u0433\u043d\u0430\u043b\u043e\u0432 \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u0430 \u2192 deterministic_competitor_fallback \u2192 monitor_queue.\n\n\u041b\u0438\u043c\u0438\u0442\u044b: \u043c\u0430\u043a\u0441\u0438\u043c\u0443\u043c 5 URL (\u0436\u0451\u0441\u0442\u043a\u0430\u044f \u043e\u0442\u0441\u0435\u0447\u043a\u0430), \u043f\u0435\u0440\u0432\u044b\u0439 \u0437\u0430\u043f\u0443\u0441\u043a \u2014 3 URL. \u041d\u0435 \u0430\u043a\u0442\u0438\u0432\u0438\u0440\u043e\u0432\u0430\u0442\u044c. \u0421\u0445\u0435\u043c\u0430 35 \u043a\u043e\u043b\u043e\u043d\u043e\u043a (\u0441 run_id, batch_index); url_registry \u2014 \u0441\u0432\u043e\u0438 10 \u043a\u043e\u043b\u043e\u043d\u043e\u043a.\n\u041e\u0436\u0438\u0434\u0430\u0435\u043c\u043e: \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0430 \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u0430 \u2192 monitor_queue; \u0434\u0443\u0431\u043b\u0438\u043a\u0430\u0442 \u2192 skipped_log.",
        "height": 420,
        "width": 540,
        "color": 4
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000002",
      "name": "Manual Start",
      "type": "n8n-nodes-base.manualTrigger",
      "typeVersion": 1,
      "position": [
        120,
        300
      ],
      "parameters": {}
    },
    {
      "id": "b4000000-0000-0000-0000-00000000000l",
      "name": "Set URL List",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        220,
        300
      ],
      "parameters": {
        "jsCode": "// OPERATOR: edit rawUrls below. MAX 5 URLs (hard cap). First run: use 3.\n// Keep ONLY example.com placeholders committed to the repo \u2014 paste real URLs locally.\nvar __agentIn = (typeof $json === 'object' && $json) ? $json : {};\nconst __exampleUrls = [\n  'https://example.com/competitor-1',\n  'https://example.com/competitor-2',\n  'https://example.com/competitor-3'\n];\nconst rawUrls = (Array.isArray(__agentIn.urls) && __agentIn.urls.length) ? __agentIn.urls : __exampleUrls;\n\nfunction pad(n){ return String(n).padStart(2, '0'); }\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nfunction moscowStamp(){var z=function(n){return String(n).padStart(2,'0');};var m=new Date(Date.now()+10800000);return m.getUTCFullYear()+z(m.getUTCMonth()+1)+z(m.getUTCDate())+'_'+z(m.getUTCHours())+z(m.getUTCMinutes())+z(m.getUTCSeconds());}\nconst now = moscowIsoNow();\nconst __stamp = moscowStamp();\nconst run_id = (__agentIn.source_run_id || __agentIn.run_id) ? String(__agentIn.source_run_id || __agentIn.run_id) : ('firecrawl_' + __stamp);\nconst agent_request_id = __agentIn.agent_request_id ? String(__agentIn.agent_request_id) : ('wf04_req_' + __stamp); // distinct request id (not the source_run_id)\n// SOURCE-REUSE-001: the caller's owner + execution-mode context. WF04 must make the SAME reuse/collect/refresh\n// decision the planner promised, so the planner's inputs travel with every item.\nconst owner_user_id = __agentIn.owner_user_id ? String(__agentIn.owner_user_id) : '';\nconst source_execution_mode = __agentIn.source_execution_mode ? String(__agentIn.source_execution_mode) : '';\nconst freshness_days = Number(__agentIn.freshness_days) || 0;\n\nconst cleaned = [];\nfor (let u of rawUrls) {\n  if (u == null) continue;\n  u = String(u).trim();\n  if (u === '') continue;\n  cleaned.push(u);\n}\nconst capped = cleaned.slice(0, 5); // HARD CAP: max 5 URLs\n\n// --- Stage C Closure Patch 2: reset run-level repair/fallback accounting (S2-D7/D8/D9/D16) ---\nconst __sd=$getWorkflowStaticData('global');\n__sd.wf04_run={ urls_received:capped.length, urls_scraped:0, primary_calls:0, primary_parse_success:0, primary_parse_failure:0, repair_calls:0, repair_success:0, repair_failure:0, deterministic_fallback:0, degraded:0, quarantined:0, snapshots_written:0, firecrawl_calls:0, claude_calls:0, reused:0, reuse_failed:0, original_snapshot_run_id:'', original_snapshot_collected_at:'', force_reprocess:(__agentIn.force_reprocess === true || __agentIn.force_reprocess === 'true'), requested_mode:source_execution_mode };\nconst items = [];\nlet idx = 0;\nfor (const u of capped) {\n  idx++;\n  items.push({ json: {\n    target_url: u,\n    source_type: 'scraped_web',\n    platform: 'website',\n    parsed_at: now,\n    source_note: 'firecrawl_url_list_manual',\n    run_id: run_id,\n    agent_request_id: agent_request_id,\n    batch_index: idx,\n    force_reprocess: (__agentIn.force_reprocess === true || __agentIn.force_reprocess === 'true'), // FORCE-REPROCESS-001: callable passes strings\n    owner_user_id: owner_user_id,\n    source_execution_mode: source_execution_mode,\n    freshness_days: freshness_days\n  }});\n}\nreturn items;"
      }
    },
    {
      "id": "b4000000-0000-0000-0000-00000000000b",
      "name": "Loop Over Items",
      "type": "n8n-nodes-base.splitInBatches",
      "typeVersion": 3,
      "position": [
        440,
        300
      ],
      "parameters": {
        "options": {}
      }
    },
    {
      "id": "b4000000-0000-0000-0000-00000000000n",
      "name": "Normalize URL for Dedup",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        660,
        300
      ],
      "parameters": {
        "jsCode": "const j = $json;\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nlet u = String(j.target_url || '').trim();\nlet normalized = u;\ntry {\n  if (u) {\n    const noFrag = u.split('#')[0];\n    const url = new URL(noFrag);\n    url.protocol = (url.protocol || '').toLowerCase();\n    url.hostname = (url.hostname || '').toLowerCase();\n    const drop = ['utm_source','utm_medium','utm_campaign','utm_term','utm_content','gclid','yclid','fbclid'];\n    for (const p of drop) url.searchParams.delete(p);\n    let path = url.pathname || '/';\n    if (path.length > 1 && path.endsWith('/')) path = path.slice(0, -1);\n    url.pathname = path;\n    normalized = url.toString();\n  }\n} catch(e) {\n  normalized = u;\n}\nreturn [{ json: {\n  target_url: u || normalized,\n  normalized_source_url: normalized,\n  source_url: normalized,\n  source_type: j.source_type || 'scraped_web',\n  platform: j.platform || 'website',\n  parsed_at: j.parsed_at || moscowIsoNow(),\n  source_note: j.source_note || 'firecrawl_url_list_manual',\n  run_id: j.run_id || '',\n  batch_index: j.batch_index || 0,\n  force_reprocess: (j.force_reprocess === true || j.force_reprocess === 'true'),\n  owner_user_id: j.owner_user_id || '',\n  source_execution_mode: j.source_execution_mode || '',\n  freshness_days: Number(j.freshness_days) || 0\n}}];"
      }
    },
    {
      "id": "b4000000-0000-0000-0000-0000000000rg",
      "name": "Registry Lookup",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        880,
        300
      ],
      "alwaysOutputData": true,
      "onError": "continueRegularOutput",
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "read",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "url_registry",
          "mode": "name"
        },
        "filtersUI": {
          "values": [
            {
              "lookupColumn": "normalized_source_url",
              "lookupValue": "={{ $('Normalize URL for Dedup').first().json.normalized_source_url }}"
            }
          ]
        },
        "options": {
          "returnAllMatches": "returnAllMatches"
        }
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "b4000000-0000-0000-0000-00000000000e",
      "name": "Evaluate Dedup",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1100,
        300
      ],
      "parameters": {
        "jsCode": "const ctx = $('Normalize URL for Dedup').first().json;\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nconst now = moscowIsoNow();\nconst key = String(ctx.normalized_source_url || '').trim();\nconst run_id = ctx.run_id || '';\nconst batch_index = ctx.batch_index || 0;\nconst parsedAt = ctx.parsed_at || now;\nconst force = (ctx.force_reprocess === true || ctx.force_reprocess === 'true');\nconst ownerId = String(ctx.owner_user_id || '');\nconst freshnessDays = Number(ctx.freshness_days) || 0;\n\n// embedded n8n/lib/source_execution_policy.js (do not edit here; edit the lib and re-run the transform)\n// source_execution_policy.js \u2014 the ONE decision for \"do we pay to collect this source again?\"\n//\n// SOURCE-EXEC-001. Before this, WF04's `url_registry` check was a PERMANENT dedup: `hit = !force && rows.some(...)`\n// with no time component. Once a URL was ever scraped it was skipped forever, so a user could never re-analyze a\n// site \u2014 and the run produced an empty bundle and a misleading \"\u0434\u0430\u043d\u043d\u044b\u0445 \u043d\u0435\u0442\" while a perfectly good saved snapshot\n// sat in the sheet. Permanent-skip is not a freshness policy; it is a leak.\n//\n// Three explicit modes:\n//   reuse   \u2014 a recent ACCEPTED snapshot exists and is still fresh. No paid collection; analyze the saved snapshot\n//             and say so, with its collection time. Collection cost $0.\n//   collect \u2014 no accepted snapshot, or the newest is older than the TTL, or every prior attempt failed. Pay once.\n//   refresh \u2014 the user explicitly asked to re-collect. Bypasses ONLY the freshness check; approval, budget, quality\n//             and global dedup all still apply, and the repeated paid collection is stated in the plan.\n//\n// A FAILED snapshot never counts as reusable and never blocks a retry \u2014 otherwise one bad scrape would poison a\n// source forever (the same permanence bug in a different costume).\n//\n// Embeddable: unique sx*-prefixed names, no cross-lib require.\n\nfunction sxStr(v) { return v == null ? '' : String(v); }\nfunction sxNum(v, d) { var n = Number(v); return isFinite(n) ? n : (d === undefined ? 0 : d); }\n\nvar SX_MODES = { REUSE: 'reuse', COLLECT: 'collect', REFRESH: 'refresh' };\n// Default freshness window. The product's monitoring/report cadence is weekly (WF25 weekly digest, WF23 monitor),\n// and WF10/WF12 already reason over a 30-day window \u2014 so a snapshot younger than 7 days is \"current\" for a\n// competitor's public positioning (offers/prices change on a weeks-to-months scale, not hourly). Operator override:\n// MS_SOURCE_FRESHNESS_DAYS.\nvar SX_DEFAULT_FRESHNESS_DAYS = 7;\n\n// A snapshot is reusable only if it actually produced accepted content.\nvar SX_REJECTED_STATUSES = ['failed', 'error', 'quarantined', 'technical_error', 'excluded', 'invalid'];\nfunction sxIsAccepted(s) {\n  if (!s) return false;\n  var st = sxStr(s.quality_status || s.status).toLowerCase();\n  if (st && SX_REJECTED_STATUSES.indexOf(st) >= 0) return false;\n  // SOURCE-REUSE-001: url_registry rows carry the verdict in processing_status (last_processing_status), not in\n  // quality_status. Checking only technical_error here let a quarantined/failed prior run count as \"accepted\".\n  var ps = sxStr(s.processing_status).toLowerCase();\n  if (ps && SX_REJECTED_STATUSES.indexOf(ps) >= 0) return false;\n  if (s.accepted === false) return false;\n  return true;\n}\nfunction sxTime(s) {\n  var t = Date.parse(sxStr((s && (s.collected_at || s.parsed_at || s.last_seen_at || s.created_at)) || ''));\n  return isFinite(t) ? t : NaN;\n}\nfunction sxNorm(u) {\n  return sxStr(u).trim().toLowerCase().replace(/^https?:\\/\\//, '').replace(/^www\\./, '').replace(/\\/+$/, '');\n}\n\n// newestAcceptedSnapshot(snapshots, source_url, owner) -> the freshest ACCEPTED snapshot for this url (or null).\n// Owner isolation: when a snapshot carries an owner, only that owner's rows are considered.\nfunction newestAcceptedSnapshot(snapshots, sourceUrl, owner) {\n  var key = sxNorm(sourceUrl);\n  var best = null, bestT = -1;\n  (Array.isArray(snapshots) ? snapshots : []).forEach(function (s) {\n    if (!s) return;\n    if (sxNorm(s.source_url || s.normalized_source_url) !== key) return;\n    if (owner && sxStr(s.owner_user_id) && sxStr(s.owner_user_id) !== sxStr(owner)) return;\n    if (!sxIsAccepted(s)) return;\n    var t = sxTime(s);\n    if (isNaN(t)) return;\n    if (t > bestT) { bestT = t; best = s; }\n  });\n  return best;\n}\n\n// decideSourceExecution(input) -> { mode, reason, force_reprocess, snapshot, snapshot_age_days, snapshot_collected_at }\n// input: { source_url, snapshots[], owner_user_id, requested_refresh, now, freshness_days }\nfunction decideSourceExecution(input) {\n  input = input || {};\n  var nowMs = input.now ? Date.parse(sxStr(input.now)) : Date.now();\n  if (!isFinite(nowMs)) nowMs = Date.now();\n  var ttlDays = sxNum(input.freshness_days, SX_DEFAULT_FRESHNESS_DAYS);\n  if (!(ttlDays > 0)) ttlDays = SX_DEFAULT_FRESHNESS_DAYS;\n\n  var snap = newestAcceptedSnapshot(input.snapshots, input.source_url, input.owner_user_id);\n  // Did we try this url before and fail? That is NOT \"never collected\" \u2014 the user deserves the real reason, and a\n  // failed attempt must never block a retry.\n  var key = sxNorm(input.source_url);\n  var triedBefore = (Array.isArray(input.snapshots) ? input.snapshots : []).some(function (s) {\n    return s && sxNorm(s.source_url || s.normalized_source_url) === key &&\n      (!input.owner_user_id || !sxStr(s.owner_user_id) || sxStr(s.owner_user_id) === sxStr(input.owner_user_id));\n  });\n  var ageDays = snap ? (nowMs - sxTime(snap)) / 86400000 : null;\n  var out = {\n    mode: SX_MODES.COLLECT, reason: triedBefore ? 'last_attempt_failed' : 'never_collected', force_reprocess: false,\n    snapshot: null, snapshot_age_days: null, snapshot_collected_at: ''\n  };\n  if (snap) {\n    out.snapshot = snap;\n    out.snapshot_age_days = Math.round(ageDays * 100) / 100;\n    out.snapshot_collected_at = sxStr(snap.collected_at || snap.parsed_at || snap.last_seen_at || snap.created_at);\n  }\n\n  // An explicit refresh wins over freshness \u2014 but ONLY over freshness. Everything else still gates it.\n  if (input.requested_refresh === true) {\n    out.mode = SX_MODES.REFRESH;\n    out.force_reprocess = true;\n    out.reason = snap ? 'explicit_refresh' : 'explicit_refresh_no_snapshot';\n    return out;\n  }\n  if (!snap) return out;                                   // collect / never_collected\n  if (ageDays > ttlDays) { out.reason = 'snapshot_stale'; return out; }  // collect / stale\n  out.mode = SX_MODES.REUSE;\n  out.reason = 'fresh_snapshot';\n  return out;\n}\n\n// Russian phrases that mean \"collect it again, now\". Cyrillic \\b/\\w do not fire in JS, so match on explicit\n// [\u0430-\u044f\u0451] boundaries. Deliberately narrow: an accidental refresh costs the user real money.\nvar SX_REFRESH_RE = /(^|[^\u0430-\u044f\u0451a-z])(\u043e\u0431\u043d\u043e\u0432\u0438(\u0442\u044c|\u0442\u0435)?|\u043f\u0435\u0440\u0435\u043e\u0431\u043d\u043e\u0432\u0438(\u0442\u044c|\u0442\u0435)?|\u043f\u0435\u0440\u0435\u0441\u043e\u0431\u0435\u0440\u0438|\u043f\u0435\u0440\u0435\u0441\u043e\u0431\u0440\u0430\u0442\u044c|\u043f\u0435\u0440\u0435\u0441\u043e\u0431\u0435\u0440\u0438\u0442\u0435|\u0437\u0430\u043d\u043e\u0432\u043e|\u043f\u043e\u0432\u0442\u043e\u0440\u0438(\u0442\u044c|\u0442\u0435)? \u0441\u0431\u043e\u0440|\u043f\u043e\u0432\u0442\u043e\u0440\u043d\u044b\u0439 \u0441\u0431\u043e\u0440|\u043f\u0440\u0438\u043d\u0443\u0434\u0438\u0442\u0435\u043b\u044c\u043d(\u043e|\u044b\u0439|\u0430\u044f)|\u0435\u0449\u0451 \u0440\u0430\u0437 \u0441\u043e\u0431\u0435\u0440\u0438|\u0435\u0449\u0435 \u0440\u0430\u0437 \u0441\u043e\u0431\u0435\u0440\u0438|\u0441\u0432\u0435\u0436(\u0438\u0435|\u0438\u0445) \u0434\u0430\u043d\u043d(\u044b\u0435|\u044b\u0445)|\u0430\u043a\u0442\u0443\u0430\u043b\u0438\u0437\u0438\u0440\u0443\u0439|\u043f\u0435\u0440\u0435\u043f\u0440\u043e\u0432\u0435\u0440\u044c)([^\u0430-\u044f\u0451a-z]|$)/i;\n// \"\u043e\u0431\u043d\u043e\u0432\u0438 \u043e\u0442\u0447\u0451\u0442\" = rebuild the report from what we have; it is NOT automatically a paid re-collection.\nvar SX_REPORT_ONLY_RE = /\u043e\u0431\u043d\u043e\u0432\u0438(\u0442\u044c|\u0442\u0435)?\\s+\u043e\u0442\u0447[\u0435\u0451]\u0442/i;\n\n// detectRefreshRequest(text) -> { requested_refresh, refresh_reason }\nfunction detectRefreshRequest(text) {\n  var t = sxStr(text);\n  if (!t) return { requested_refresh: false, refresh_reason: '' };\n  if (SX_REPORT_ONLY_RE.test(t) && !/\u0434\u0430\u043d\u043d|\u0441\u0431\u043e\u0440|\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a|\u0441\u0430\u0439\u0442/i.test(t)) {\n    return { requested_refresh: false, refresh_reason: 'report_rebuild_only' };\n  }\n  if (SX_REFRESH_RE.test(t)) return { requested_refresh: true, refresh_reason: 'user_requested_refresh' };\n  return { requested_refresh: false, refresh_reason: '' };\n}\n// --- end embedded source_execution_policy ---\n\n// SOURCE-REUSE-001. The old check here was a PERMANENT dedup: any registry hit emitted only a skipped_log row and\n// ZERO data, so the planner could promise \"reuse\" while the executor delivered \"no_sources\" (live: WF04 exec 948 ->\n// WF10 exec 951 rows_after_isolation=0 for a site scraped 25 minutes earlier). Planning and execution now share the\n// ONE canonical decision: decideSourceExecution (reuse | collect | refresh).\nlet rows = [];\ntry { rows = $('Registry Lookup').all().map(i => i.json); } catch(e) { rows = []; }\n// url_registry rows -> the policy's snapshot shape (the registry's last_processing_status IS the snapshot status).\nconst snapshots = rows.filter(r => r).map(function (r) { return {\n  source_url: r.normalized_source_url || r.source_url,\n  normalized_source_url: r.normalized_source_url,\n  last_seen_at: r.last_seen_at || r.first_seen_at,\n  processing_status: r.last_processing_status,\n  run_id: r.run_id,\n  last_route: r.last_route,\n  owner_user_id: r.owner_user_id // absent on today's registry rows; enforced whenever present\n};});\nconst decision = decideSourceExecution({ source_url: key, snapshots: snapshots, owner_user_id: ownerId, requested_refresh: force, freshness_days: freshnessDays, now: now });\n\n// Reuse is only real if the original run left something WF10 can actually consume: a row in one of the three data\n// queues, from a run whose source_health verdict is still eligible. Anything else -> collect again (fail closed:\n// a blocked/quarantined/unscored original never silently becomes \"saved data\").\nconst REUSABLE_ROUTES = ['monitor_queue', 'content_queue', 'review_queue'];\nlet originalHealth = null;\nif (decision.mode === 'reuse') {\n  const origRoute = String((decision.snapshot && decision.snapshot.last_route) || '');\n  if (REUSABLE_ROUTES.indexOf(origRoute) < 0) {\n    decision.mode = 'collect';\n    decision.reason = 'previous_run_not_reusable_route';\n  } else {\n    let healthRows = [];\n    try { healthRows = $('Source Health Lookup').all().map(i => i.json); } catch(e) { healthRows = []; }\n    const orig = String(decision.snapshot.run_id || '');\n    const BAD_HEALTH = ['quarantined', 'failed', 'error', 'invalid', 'excluded'];\n    healthRows.forEach(function (h) {\n      if (!h || String(h.source_run_id || '') !== orig) return;\n      const t = Date.parse(String(h.evaluated_at || '')) || 0;\n      if (!originalHealth || t >= (Date.parse(String(originalHealth.evaluated_at || '')) || 0)) originalHealth = h;\n    });\n    const hOk = originalHealth\n      && BAD_HEALTH.indexOf(String(originalHealth.quality_status || '').toLowerCase()) < 0\n      && String(originalHealth.report_eligible).toLowerCase() !== 'false';\n    if (!hOk) {\n      decision.mode = 'collect';\n      decision.reason = originalHealth ? 'original_run_not_eligible' : 'original_run_unscored';\n      originalHealth = null;\n    }\n  }\n}\n\nif (decision.mode === SX_MODES.REUSE) {\n  return [{ json: {\n    sx_mode: 'reuse',\n    sx_reason: decision.reason,\n    target_url: ctx.target_url || key,\n    normalized_source_url: key,\n    source_url: key,\n    source_type: ctx.source_type || 'scraped_web',\n    platform: ctx.platform || 'website',\n    parsed_at: parsedAt,\n    run_id: run_id,\n    batch_index: batch_index,\n    owner_user_id: ownerId,\n    reuse_route: String(decision.snapshot.last_route),\n    original_run_id: String(decision.snapshot.run_id || ''),\n    original_collected_at: String(decision.snapshot_collected_at || ''),\n    original_snapshot_age_days: decision.snapshot_age_days,\n    original_health_row: originalHealth\n  }}];\n}\n\n// collect / refresh -> Firecrawl branch (same record shape as before, now with the typed decision attached).\nreturn [{ json: {\n  sx_mode: decision.mode,\n  sx_reason: decision.reason,\n  target_url: ctx.target_url || key,\n  normalized_source_url: key,\n  source_url: key,\n  source_type: ctx.source_type || 'scraped_web',\n  platform: ctx.platform || 'website',\n  parsed_at: parsedAt,\n  run_id: run_id,\n  batch_index: batch_index\n}}];"
      }
    },
    {
      "id": "b4000000-0000-0000-0000-00000000000i",
      "name": "IF Reuse?",
      "type": "n8n-nodes-base.if",
      "typeVersion": 2,
      "position": [
        1320,
        300
      ],
      "parameters": {
        "conditions": {
          "options": {
            "caseSensitive": true,
            "leftValue": "",
            "typeValidation": "loose"
          },
          "conditions": [
            {
              "id": "b4-if-dup-01",
              "leftValue": "={{ $json.sx_mode }}",
              "rightValue": "reuse",
              "operator": {
                "type": "string",
                "operation": "equals"
              }
            }
          ],
          "combinator": "and"
        }
      }
    },
    {
      "id": "fc000000-0000-0000-0000-000000000002",
      "name": "Build Firecrawl Request",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1540,
        300
      ],
      "parameters": {
        "jsCode": "const url = $json.target_url || '';\nreturn [{ json: {\n  url: url,\n  formats: ['markdown'],\n  onlyMainContent: true,\n  onlyCleanContent: false,\n  removeBase64Images: true,\n  blockAds: true,\n  timeout: 60000,\n  storeInCache: true\n}}];"
      }
    },
    {
      "id": "fc000000-0000-0000-0000-000000000003",
      "name": "Firecrawl Scrape API",
      "type": "n8n-nodes-base.httpRequest",
      "typeVersion": 4.2,
      "position": [
        1760,
        300
      ],
      "onError": "continueRegularOutput",
      "parameters": {
        "method": "POST",
        "url": "https://api.firecrawl.dev/v2/scrape",
        "authentication": "predefinedCredentialType",
        "nodeCredentialType": "httpHeaderAuth",
        "sendHeaders": true,
        "headerParameters": {
          "parameters": [
            {
              "name": "Content-Type",
              "value": "application/json"
            }
          ]
        },
        "sendBody": true,
        "contentType": "raw",
        "rawContentType": "application/json",
        "body": "={{ JSON.stringify($json) }}",
        "options": {}
      },
      "credentials": {
        "httpHeaderAuth": {
          "name": "<your credential>"
        }
      }
    },
    {
      "id": "fc000000-0000-0000-0000-000000000004",
      "name": "Normalize Firecrawl Output",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1980,
        300
      ],
      "parameters": {
        "jsCode": "// embedded n8n/lib/source_access.js (drift-proof; test asserts equality)\n// source_access.js \u2014 BLOCK-HONESTY-001. Decide whether we actually REACHED a source, BEFORE anyone asks whether the\n// business is relevant.\n//\n// The bug this exists to kill (live: carmoney.ru, WF04 exec 894): Firecrawl returned a 772-char access-restriction\n// page. It was longer than the 80-char \"meaningful content\" floor, so it flowed into business classification, which\n// \u2014 correctly, given the text it saw \u2014 said `entity_type=irrelevant / business_skip`. The user was then told\n// \u00abcarmoney.ru \u2014 \u043f\u0440\u043e\u0432\u0435\u0440\u0435\u043d, \u043d\u043e\u0432\u044b\u0445 \u0440\u0435\u043b\u0435\u0432\u0430\u043d\u0442\u043d\u044b\u0445 \u0444\u0430\u043a\u0442\u043e\u0432 \u043d\u0435 \u043d\u0430\u0439\u0434\u0435\u043d\u043e\u00bb and advised to widen filters. Every part of that is\n// false: we never saw the site, the company may be highly relevant, and no filter change can fix an IP block.\n//\n// \"We could not read the page\" and \"we read the page and it is not a competitor\" are DIFFERENT facts with different\n// user messages, different next actions, and different data consequences. Access is decided first, from the\n// transport + the page's own text; relevance is only asked once access is `accessible_content`.\n//\n// A non-accessible outcome may NEVER become a competitor snapshot, overwrite a good snapshot, or feed Claude.\n//\n// Embeddable: unique sa*-prefixed names, no cross-lib require.\n\nfunction saStr(v) { return v == null ? '' : String(v); }\nfunction saNum(v, d) { var n = Number(v); return isFinite(n) ? n : (d === undefined ? 0 : d); }\n\nvar SA_OUTCOMES = {\n  ACCESSIBLE: 'accessible_content',\n  BLOCKED_WAF: 'blocked_by_waf',\n  ACCESS_DENIED: 'robots_or_access_denied',\n  PROVIDER_FAILURE: 'provider_failure',\n  TIMEOUT: 'timeout',\n  EMPTY: 'empty_response',\n  UNSUPPORTED: 'unsupported_content',\n  IRRELEVANT: 'valid_but_irrelevant'   // set DOWNSTREAM, only after ACCESSIBLE \u2014 never inferred here\n};\n// Only these mean \"we hold real page content\".\nfunction saIsAccessible(o) { return o === SA_OUTCOMES.ACCESSIBLE; }\n// These are transport/access problems: the business is unjudged, so the user must never be told it is irrelevant.\nfunction saIsAccessFailure(o) {\n  return [SA_OUTCOMES.BLOCKED_WAF, SA_OUTCOMES.ACCESS_DENIED, SA_OUTCOMES.PROVIDER_FAILURE,\n    SA_OUTCOMES.TIMEOUT, SA_OUTCOMES.EMPTY, SA_OUTCOMES.UNSUPPORTED].indexOf(o) >= 0;\n}\n// Worth another attempt later (a block/timeout may lift); an unsupported content type will not fix itself.\nfunction saIsRetryable(o) {\n  return [SA_OUTCOMES.BLOCKED_WAF, SA_OUTCOMES.PROVIDER_FAILURE, SA_OUTCOMES.TIMEOUT, SA_OUTCOMES.EMPTY].indexOf(o) >= 0;\n}\n\n// STRONG signatures: unambiguous challenge/block boilerplate. A real commercial page does not say these about itself.\nvar SA_STRONG_WAF = [\n  'just a moment...', 'checking your browser before accessing', 'attention required! | cloudflare',\n  'cloudflare ray id', 'enable javascript and cookies to continue', 'verify you are human',\n  'ddos protection by cloudflare', 'performance & security by cloudflare', 'error 1020', 'error code 1020',\n  'ray id:', 'sorry, you have been blocked', 'why have i been blocked', 'incapsula incident id',\n  'request unsuccessful. incapsula', 'access to this page has been denied', 'pardon our interruption',\n  'are you a robot', '\u043f\u043e\u0434\u0442\u0432\u0435\u0440\u0434\u0438\u0442\u0435, \u0447\u0442\u043e \u0432\u044b \u043d\u0435 \u0440\u043e\u0431\u043e\u0442', '\u043f\u0440\u043e\u0432\u0435\u0440\u043a\u0430 \u0431\u0440\u0430\u0443\u0437\u0435\u0440\u0430', '\u0434\u043e\u0441\u0442\u0443\u043f \u043e\u0433\u0440\u0430\u043d\u0438\u0447\u0435\u043d'\n];\nvar SA_STRONG_DENIED = [\n  '403 forbidden', 'http error 403', 'access denied', 'you don\\'t have permission to access',\n  'you do not have permission to access', 'authorization required', '401 unauthorized',\n  '\u0434\u043e\u0441\u0442\u0443\u043f \u0437\u0430\u043f\u0440\u0435\u0449\u0451\u043d', '\u0434\u043e\u0441\u0442\u0443\u043f \u0437\u0430\u043f\u0440\u0435\u0449\u0435\u043d', '\u043d\u0435\u0442 \u0434\u043e\u0441\u0442\u0443\u043f\u0430 \u043a \u044d\u0442\u043e\u0439 \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0435'\n];\n// The live carmoney.ru family: an IP/geo restriction notice. Two independent phrases must co-occur, so a page that\n// merely mentions \"VPN\" as a product never trips it.\nvar SA_GEOBLOCK_PAIRS = [\n  ['is not available', 'restricted access from your current network'],\n  ['is not available', 'block access from specific countries'],\n  ['restricted access from your current network', 'enabled a vpn'],\n  ['website owner has restricted access', 'ip addresses'],\n  ['\u043d\u0435\u0434\u043e\u0441\u0442\u0443\u043f\u0435\u043d', '\u043e\u0433\u0440\u0430\u043d\u0438\u0447\u0438\u043b \u0434\u043e\u0441\u0442\u0443\u043f']\n];\n// WEAK signals: only meaningful on a SHORT page with no commercial content of its own.\nvar SA_WEAK = ['captcha', 'recaptcha', 'hcaptcha', 'challenge-platform', 'cf-browser-verification',\n  'security check', 'bot detection', 'unusual traffic', 'rate limited', 'too many requests'];\n\n// A page that talks about lending/pricing/contacts is a real page, whatever boilerplate it also contains.\nvar SA_BUSINESS_TERMS = ['\u043a\u0440\u0435\u0434\u0438\u0442', '\u0437\u0430\u0439\u043c', '\u0437\u0430\u043b\u043e\u0433', '\u043f\u0442\u0441', '\u0441\u0442\u0430\u0432\u043a\u0430', '\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441', '\u0438\u043f\u043e\u0442\u0435\u043a', '\u043e\u0434\u043e\u0431\u0440\u0435\u043d', '\u0437\u0430\u044f\u0432\u043a',\n  '\u0442\u0430\u0440\u0438\u0444', '\u0443\u0441\u043b\u0443\u0433', '\u043e\u0444\u043e\u0440\u043c\u0438\u0442\u044c', '\u0440\u0443\u0431', '\u043f\u0440\u043e\u0446\u0435\u043d\u0442', '\u043e\u0444\u0438\u0441', 'loan', 'credit', 'rate', 'apply'];\n\nfunction saHasBusiness(low) {\n  var n = 0;\n  for (var i = 0; i < SA_BUSINESS_TERMS.length; i++) if (low.indexOf(SA_BUSINESS_TERMS[i]) >= 0) n++;\n  return n;\n}\nfunction saAnyHit(low, list) {\n  for (var i = 0; i < list.length; i++) if (low.indexOf(list[i]) >= 0) return list[i];\n  return '';\n}\nfunction saPairHit(low, pairs) {\n  for (var i = 0; i < pairs.length; i++) {\n    if (low.indexOf(pairs[i][0]) >= 0 && low.indexOf(pairs[i][1]) >= 0) return pairs[i].join(' + ');\n  }\n  return '';\n}\n\n// classifySourceAccess(input) -> { outcome, reason, signature, retryable, access_failure, meaningful_chars, business_terms }\n// input: { status, body_text, error, error_category, content_type, url }\n// Order matters: transport verdicts first (they are authoritative), then the page's own text.\nfunction classifySourceAccess(input) {\n  input = input || {};\n  var status = saNum(input.status, NaN);\n  var text = saStr(input.body_text);\n  var low = text.toLowerCase();\n  var meaningful = text.replace(/[#>*_|[\\]()]/g, ' ').replace(/\\s+/g, ' ').trim();\n  var biz = saHasBusiness(low);\n  var out = function (outcome, reason, signature) {\n    return {\n      outcome: outcome, reason: reason, signature: saStr(signature),\n      retryable: saIsRetryable(outcome), access_failure: saIsAccessFailure(outcome),\n      meaningful_chars: meaningful.length, business_terms: biz\n    };\n  };\n\n  // 1. transport-level truth\n  var ec = saStr(input.error_category).toLowerCase();\n  if (ec === 'timeout' || /timeout|etimedout|timed out/i.test(saStr(input.error))) return out(SA_OUTCOMES.TIMEOUT, 'provider_timeout', ec || 'timeout');\n  if (input.error) return out(SA_OUTCOMES.PROVIDER_FAILURE, 'provider_error', saStr(input.error).slice(0, 60));\n  if (isFinite(status)) {\n    if (status === 401 || status === 403) return out(SA_OUTCOMES.ACCESS_DENIED, 'http_' + status, 'http_' + status);\n    if (status === 429) return out(SA_OUTCOMES.BLOCKED_WAF, 'http_429_rate_limited', 'http_429');\n    if (status >= 500) return out(SA_OUTCOMES.PROVIDER_FAILURE, 'http_' + status, 'http_' + status);\n    if (status >= 400) return out(SA_OUTCOMES.PROVIDER_FAILURE, 'http_' + status, 'http_' + status);\n  }\n  // 2. content type we cannot analyze\n  var ct = saStr(input.content_type).toLowerCase();\n  if (ct && !/text\\/html|text\\/plain|application\\/xhtml|markdown|application\\/json/.test(ct)) {\n    return out(SA_OUTCOMES.UNSUPPORTED, 'unsupported_content_type', ct.slice(0, 40));\n  }\n  // 3. nothing came back\n  if (!meaningful) return out(SA_OUTCOMES.EMPTY, 'empty_body', '');\n\n  // 4. the page's own text says we were blocked. Strong signatures win regardless of length: real commercial pages\n  //    do not describe themselves as a browser challenge.\n  var s = saAnyHit(low, SA_STRONG_WAF);\n  if (s) return out(SA_OUTCOMES.BLOCKED_WAF, 'waf_challenge_page', s);\n  var g = saPairHit(low, SA_GEOBLOCK_PAIRS);\n  if (g) return out(SA_OUTCOMES.BLOCKED_WAF, 'network_or_geo_restriction', g);\n  var d = saAnyHit(low, SA_STRONG_DENIED);\n  if (d) return out(SA_OUTCOMES.ACCESS_DENIED, 'access_denied_page', d);\n\n  // 5. weak signals: only on a short page that carries no commercial content of its own \u2014 otherwise a lender that\n  //    happens to mention \"captcha\" in its FAQ would be wrongly reported as blocked.\n  if (meaningful.length < 1200 && biz === 0) {\n    var w = saAnyHit(low, SA_WEAK);\n    if (w) return out(SA_OUTCOMES.BLOCKED_WAF, 'short_page_block_signature', w);\n  }\n  // 6. too little to analyze (mirrors WF04's existing floor)\n  if (meaningful.length < 80) return out(SA_OUTCOMES.EMPTY, 'body_too_short', String(meaningful.length) + ' chars');\n\n  return out(SA_OUTCOMES.ACCESSIBLE, 'content_ok', '');\n}\n\n// The ONE user-facing Russian sentence per access failure. Never a status code, provider name, or raw page text.\n// Every one states explicitly that the business was NOT judged \u2014 that is the whole point of BLOCK-HONESTY-001.\nvar SA_USER_RU = {\n  blocked_by_waf: '\u0441\u0430\u0439\u0442 \u0432\u0435\u0440\u043d\u0443\u043b \u0437\u0430\u0449\u0438\u0442\u043d\u0443\u044e \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0443 \u0438 \u043d\u0435 \u043e\u0442\u0434\u0430\u043b \u0441\u043e\u0434\u0435\u0440\u0436\u0438\u043c\u043e\u0435 \u2014 \u043f\u0440\u043e\u0447\u0438\u0442\u0430\u0442\u044c \u0435\u0433\u043e \u043d\u0435 \u0443\u0434\u0430\u043b\u043e\u0441\u044c',\n  robots_or_access_denied: '\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a \u0437\u0430\u043a\u0440\u044b\u043b \u0434\u043e\u0441\u0442\u0443\u043f \u043a \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0435',\n  provider_failure: '\u0441\u0435\u0440\u0432\u0438\u0441 \u0441\u0431\u043e\u0440\u0430 \u0434\u0430\u043d\u043d\u044b\u0445 \u043d\u0435 \u0441\u043c\u043e\u0433 \u043f\u043e\u043b\u0443\u0447\u0438\u0442\u044c \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0443',\n  timeout: '\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a \u043d\u0435 \u043e\u0442\u0432\u0435\u0442\u0438\u043b \u0432\u043e\u0432\u0440\u0435\u043c\u044f',\n  empty_response: '\u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0430 \u043e\u0442\u043a\u0440\u044b\u043b\u0430\u0441\u044c \u043f\u0443\u0441\u0442\u043e\u0439 \u2014 \u0441\u043e\u0434\u0435\u0440\u0436\u0438\u043c\u043e\u0433\u043e \u0434\u043b\u044f \u0430\u043d\u0430\u043b\u0438\u0437\u0430 \u043d\u0435\u0442',\n  unsupported_content: '\u043f\u043e \u044d\u0442\u043e\u043c\u0443 \u0430\u0434\u0440\u0435\u0441\u0443 \u043d\u0435 \u0442\u0435\u043a\u0441\u0442\u043e\u0432\u0430\u044f \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0430 \u2014 \u0430\u043d\u0430\u043b\u0438\u0437\u0438\u0440\u043e\u0432\u0430\u0442\u044c \u043d\u0435\u0447\u0435\u0433\u043e'\n};\nfunction saUserMessageRu(outcome) { return SA_USER_RU[saStr(outcome)] || '\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a \u0441\u0435\u0439\u0447\u0430\u0441 \u043d\u0435\u0434\u043e\u0441\u0442\u0443\u043f\u0435\u043d'; }\n\n// Cause-specific next actions (\u00a77). Never \"\u0440\u0430\u0441\u0448\u0438\u0440\u044c\u0442\u0435 \u0444\u0438\u043b\u044c\u0442\u0440\u044b\" for an access failure \u2014 no filter reaches a blocked page.\nvar SA_NEXT_RU = {\n  blocked_by_waf: ['\u043f\u043e\u0432\u0442\u043e\u0440\u0438\u0442\u044c \u043f\u043e\u043f\u044b\u0442\u043a\u0443 \u043f\u043e\u0437\u0436\u0435', '\u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u044c \u043f\u043e\u0441\u043b\u0435\u0434\u043d\u0438\u0439 \u0441\u043e\u0445\u0440\u0430\u043d\u0451\u043d\u043d\u044b\u0439 \u0441\u043d\u0438\u043c\u043e\u043a', '\u043f\u0440\u043e\u0432\u0435\u0440\u0438\u0442\u044c \u0434\u0440\u0443\u0433\u043e\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a'],\n  robots_or_access_denied: ['\u043f\u0440\u043e\u0432\u0435\u0440\u0438\u0442\u044c \u0434\u0440\u0443\u0433\u043e\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a', '\u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u044c \u043f\u043e\u0441\u043b\u0435\u0434\u043d\u0438\u0439 \u0441\u043e\u0445\u0440\u0430\u043d\u0451\u043d\u043d\u044b\u0439 \u0441\u043d\u0438\u043c\u043e\u043a'],\n  provider_failure: ['\u043f\u043e\u0432\u0442\u043e\u0440\u0438\u0442\u044c \u043f\u043e\u043f\u044b\u0442\u043a\u0443', '\u043f\u0440\u043e\u0432\u0435\u0440\u0438\u0442\u044c \u0434\u0440\u0443\u0433\u043e\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a'],\n  timeout: ['\u043f\u043e\u0432\u0442\u043e\u0440\u0438\u0442\u044c \u043f\u043e\u043f\u044b\u0442\u043a\u0443', '\u043f\u0440\u043e\u0432\u0435\u0440\u0438\u0442\u044c \u0434\u0440\u0443\u0433\u043e\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a'],\n  empty_response: ['\u0443\u043a\u0430\u0437\u0430\u0442\u044c \u043a\u043e\u043d\u043a\u0440\u0435\u0442\u043d\u0443\u044e \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0443 \u0443\u0441\u043b\u0443\u0433\u0438', '\u043f\u0440\u043e\u0432\u0435\u0440\u0438\u0442\u044c \u0434\u0440\u0443\u0433\u043e\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a'],\n  unsupported_content: ['\u0443\u043a\u0430\u0437\u0430\u0442\u044c \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0443 \u0441 \u0442\u0435\u043a\u0441\u0442\u043e\u043c (\u043d\u0430\u043f\u0440\u0438\u043c\u0435\u0440, \u0440\u0430\u0437\u0434\u0435\u043b \u0443\u0441\u043b\u0443\u0433)']\n};\nfunction saNextActionsRu(outcome, opts) {\n  var list = (SA_NEXT_RU[saStr(outcome)] || ['\u043f\u043e\u0432\u0442\u043e\u0440\u0438\u0442\u044c \u043f\u043e\u043f\u044b\u0442\u043a\u0443 \u043f\u043e\u0437\u0436\u0435']).slice();\n  // Only offer the saved snapshot when one actually exists.\n  if (!(opts && opts.has_snapshot)) list = list.filter(function (x) { return x.indexOf('\u0441\u043e\u0445\u0440\u0430\u043d\u0451\u043d\u043d') < 0; });\n  return list.slice(0, 3);\n}\n// --- end embedded source_access ---\n\nconst resp = $json;\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nconst ctx = $('Evaluate Dedup').first().json;\nconst targetUrl = ctx.target_url || ctx.source_url || '';\nconst sourceUrlBase = ctx.source_url || targetUrl;\nconst run_id = ctx.run_id || '';\nconst batch_index = ctx.batch_index || 0;\nconst parsedAt = ctx.parsed_at || moscowIsoNow();\nconst now = moscowIsoNow();\nfunction cap(s, n) { return (s == null ? '' : String(s)).substring(0, n); }\n\nfunction technicalErrorRow(errSummary, preview) {\n  return [{ json: {\n    created_at: now, source_type: 'scraped_web', platform: 'website', source_url: sourceUrlBase, parsed_at: parsedAt,\n    published_at: '', freshness_status: 'unknown', entity_type: 'irrelevant', company_name: '', profile_name: '',\n    profile_url: '', region: '', service_type: 'unknown', offer_text: '', terms: '', contact_public: '',\n    text_context: '', detected_need: '', competitor_strength: 1, lead_signal_score: 1, content_idea_score: 1,\n    quality_score: 1, reason: '', recommended_action: 'ignore', status: 'skipped',\n    processing_status: 'technical_error', parse_method: 'firecrawl_error',\n    parse_error: cap('Firecrawl scrape failed: ' + errSummary, 800),\n    raw_response_preview: cap(preview, 500), route: 'technical_errors', needs_manual_review: true,\n    repair_used: false, repair_status: '', run_id: run_id, batch_index: batch_index\n  }}];\n}\n\nconst apiError = (resp == null) || resp.error || resp.success === false || (resp.code && resp.code >= 400) || (resp.statusCode && resp.statusCode >= 400);\nif (apiError) {\n  const summary = cap(JSON.stringify((resp && (resp.error || resp.message || resp.code || resp.statusCode)) || 'unknown error'), 300);\n  return technicalErrorRow(summary, cap(JSON.stringify(resp), 500));\n}\n\nconst data = resp.data || resp;\nlet markdown = (resp.data && resp.data.markdown) || resp.markdown || (data && data.markdown) || (resp.data && resp.data.data && resp.data.data.markdown) || '';\nconst metadata = (resp.data && resp.data.metadata) || resp.metadata || (data && data.metadata) || {};\nmarkdown = String(markdown).replace(/\\r/g, '');\n\n// Clean: drop image lines and inline image/svg data fragments\nconst cleanedLines = [];\nfor (const line of markdown.split('\\n')) {\n  const t = line.trim();\n  if (t.startsWith('![')) continue;\n  if (/data:image\\/svg\\+xml/i.test(line)) continue;\n  if (/data:image\\/[a-z]+;base64/i.test(line)) continue;\n  cleanedLines.push(line);\n}\nlet cleaned = cleanedLines.join('\\n').replace(/\\n{3,}/g, '\\n\\n').trim();\n\nconst meaningful = cleaned.replace(/[#>*_|]/g, '').replace(/\\s+/g, ' ').trim();\n// BLOCK-HONESTY-001: decide whether we REACHED the page BEFORE anyone asks whether the business is relevant.\n// Live (carmoney.ru, exec 894): a 772-char access-restriction page cleared the 80-char floor, reached Claude, and\n// came back \"irrelevant\" \u2014 so the user was told the site was checked and found irrelevant. We never saw the site.\n// An access failure is NOT a business verdict: it must never become a competitor snapshot, never overwrite a good\n// snapshot, never reach Claude (that also saves the call), and never be reported as \"\u043f\u0440\u043e\u0432\u0435\u0440\u0435\u043d\".\nconst __access = classifySourceAccess({ status: 200, body_text: cleaned, content_type: 'text/html' });\nif (__access.access_failure) {\n  const __row = technicalErrorRow('source not accessible: ' + __access.outcome + ' (' + __access.reason + ')', cleaned)[0];\n  // entity_type stays 'unknown' \u2014 NOT 'irrelevant'. We never judged the business.\n  __row.json.entity_type = 'unknown';\n  __row.json.parse_method = 'source_' + __access.outcome;\n  __row.json.access_outcome = __access.outcome;\n  __row.json.access_retryable = __access.retryable === true;\n  __row.json.reason = saUserMessageRu(__access.outcome);\n  __row.json.recommended_action = 'ignore';\n  return [__row];\n}\nif (!cleaned || meaningful.length < 80) {\n  return technicalErrorRow('scrape succeeded but markdown empty/unusable (' + meaningful.length + ' meaningful chars)', cap(cleaned || JSON.stringify(metadata), 500));\n}\n\n// Commercially relevant lines first (stable order within groups), then cap at 3500\nconst relevantTerms = ['\u043a\u0440\u0435\u0434\u0438\u0442','\u0437\u0430\u0439\u043c','\u0437\u0430\u043b\u043e\u0433','\u043f\u0442\u0441','\u0430\u0432\u0442\u043e','\u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b','\u043d\u0435\u0434\u0432\u0438\u0436','\u043a\u0432\u0430\u0440\u0442\u0438\u0440','\u0434\u043e\u043c','\u0437\u0435\u043c\u043b','\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441','\u0438\u043f\u043e\u0442\u0435\u043a','\u0441\u0442\u0430\u0432\u043a\u0430','\u0441\u0443\u043c\u043c\u0430','\u043e\u0434\u043e\u0431\u0440\u0435\u043d','\u043f\u043b\u043e\u0445\u0430\u044f \u043a\u0440\u0435\u0434\u0438\u0442\u043d','\u043f\u0440\u043e\u0441\u0440\u043e\u0447','\u0442\u0435\u043b\u0435\u0444\u043e\u043d','\u043c\u043e\u0441\u043a\u0432\u0430','\u0440\u0443\u0431','%'];\nconst relevant = [];\nconst rest = [];\nfor (const line of cleaned.split('\\n')) {\n  const low = line.toLowerCase();\n  if (relevantTerms.some(t => low.includes(t))) relevant.push(line); else rest.push(line);\n}\nlet prioritized = (relevant.length > 0) ? (relevant.join('\\n') + '\\n' + rest.join('\\n')).replace(/\\n{3,}/g, '\\n\\n').trim() : cleaned;\n\n// Placeholder / parking / domain-not-connected pre-filter (DEC-054): skip BEFORE Claude cost.\n// Strong phrases skip unconditionally; bare 'coming soon' only when the page has no business content.\nconst placeholderHay = (cleaned + ' ' + (metadata.title || '') + ' ' + (metadata.description || '')).toLowerCase();\nconst strongPlaceholder = ['domain is not connected','domain not connected','not connected to any website','wix domain not connected','this domain is not connected','parking page','\u0441\u0430\u0439\u0442 \u043d\u0435 \u043f\u043e\u0434\u043a\u043b\u044e\u0447\u0435\u043d','\u0434\u043e\u043c\u0435\u043d \u043d\u0435 \u043f\u043e\u0434\u043a\u043b\u044e\u0447\u0435\u043d','\u0437\u0430\u0433\u043b\u0443\u0448\u043a\u0430 \u0441\u0430\u0439\u0442\u0430'];\nconst isStrongPlaceholder = strongPlaceholder.some(s => placeholderHay.includes(s));\nconst isComingSoonEmpty = placeholderHay.includes('coming soon') && relevant.length === 0;\nif (isStrongPlaceholder || isComingSoonEmpty) {\n  return [{ json: {\n    created_at: now, source_type: 'scraped_web', platform: 'website', source_url: sourceUrlBase, parsed_at: parsedAt,\n    published_at: '', freshness_status: 'unknown', entity_type: 'irrelevant', company_name: '', profile_name: '',\n    profile_url: '', region: '', service_type: 'unknown', offer_text: '', terms: '', contact_public: '',\n    text_context: '', detected_need: '', competitor_strength: 1, lead_signal_score: 1, content_idea_score: 1,\n    quality_score: 1, reason: 'Firecrawl returned a placeholder/domain-not-connected page; skipped before Claude cost.',\n    recommended_action: 'ignore', status: 'skipped', processing_status: 'business_skip',\n    parse_method: 'firecrawl_placeholder_prefilter', parse_error: '',\n    raw_response_preview: cap(cleaned, 500), route: 'skipped_log', needs_manual_review: false,\n    repair_used: false, repair_status: '', run_id: run_id, batch_index: batch_index\n  }}];\n}\n\nconst sourceUrl = metadata.sourceURL || metadata.url || sourceUrlBase;\nconst title = metadata.title || '';\nconst description = metadata.description || '';\n\nreturn [{ json: {\n  route: '', source_type: 'scraped_web', platform: 'website', source_url: sourceUrl, profile_url: '',\n  published_at: '', parsed_at: parsedAt, text_context: cap(prioritized, 3500),\n  page_title: cap(title, 300), page_description: cap(description, 500),\n  run_id: run_id, batch_index: batch_index\n}}];"
      }
    },
    {
      "id": "fc000000-0000-0000-0000-000000000005",
      "name": "IF Firecrawl Normalized OK?",
      "type": "n8n-nodes-base.if",
      "typeVersion": 2,
      "position": [
        2200,
        300
      ],
      "parameters": {
        "conditions": {
          "options": {
            "caseSensitive": true,
            "leftValue": "",
            "typeValidation": "loose"
          },
          "conditions": [
            {
              "id": "fc-if-01",
              "leftValue": "={{ $json.route }}",
              "rightValue": "",
              "operator": {
                "type": "string",
                "operation": "empty",
                "singleValue": true
              }
            }
          ],
          "combinator": "and"
        }
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000005",
      "name": "Build Primary Claude Request",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        2420,
        140
      ],
      "parameters": {
        "jsCode": "const record = $json;\n\nconst systemPrompt = `You are Marketing Scout Agent v2 -- a market intelligence analyst for a secured lending business in Moscow and Moscow Oblast, Russia.\n\nFor every record ask: What does this mean for the operator's business and what should they do? Reason like a business owner.\n\nANALYSIS PRIORITY ORDER:\n1. LEAD SIGNAL first -- potential client needing secured loan? Moscow/MO? PTS/auto/real estate collateral? Urgency? Contactable?\n2. COMPETITOR second -- active secured lending business Moscow/MO? Threat level?\n3. CONTENT IDEA third -- client fear/objection/knowledge gap for secured lending?\n4. IRRELEVANT if none apply.\n\nIDEAL CLIENT: Car owner (PTS clean title), Moscow/MO, needs cash urgently, bank-rejected, 50k-500k RUB.\n\nHigh-urgency signals (raise lead_signal_score): \u0441\u0440\u043e\u0447\u043d\u043e, \u0441\u0435\u0433\u043e\u0434\u043d\u044f, \u0431\u0430\u043d\u043a\u0438 \u043e\u0442\u043a\u0430\u0437\u0430\u043b\u0438, \u043d\u0435 \u0434\u0430\u044e\u0442 \u043a\u0440\u0435\u0434\u0438\u0442, \u0438\u0441\u043f\u043e\u0440\u0447\u0435\u043d\u0430 \u043a\u0440\u0435\u0434\u0438\u0442\u043d\u0430\u044f \u0438\u0441\u0442\u043e\u0440\u0438\u044f + specific amount + collateral type.\n\nREGION RULES:\n- Moscow/MO explicit: lead_signal_score 60-100\n- Region ambiguous: eligible up to 55\n- Another city/region: lead_signal_score capped at 40\nCompetitors: Moscow/MO or national coverage -> score normally; other region only -> cap competitor_strength at 50.\n\nlead_signal_score calibration:\n- 85-100: fit + urgency + readiness + Moscow/MO confirmed\n- 70-84: strong fit+urgency, region confirmed, readiness partial\n- 55-69: clear product fit, one signal confirmed, region present\n- 35-54: intent apparent, fit unclear or region outside MO\n- 1-34: no real lead signal\n\nrecommended_action=contact requires lead_signal_score>=70.\n\ncompetitor_strength calibration:\n- 85-100: fresh (<=30d), Moscow/MO, stated rate, same-day, bad-credit accepted, contactable\n- 65-84: active professional, Moscow/MO confirmed, rate absent or one signal missing\n- 45-64: present but older content or uncertain coverage\n- 25-44: weak - stale or different region\n- 1: not a competitor\n\nSKIP RULES -- return status=skipped, quality_score=1, all scores=1 when:\n- Fewer than 40 meaningful chars\n- Pure navigation boilerplate\n- No connection to financial services\n- published_at >180 days before parsed_at with no fresh signals\n\nREASON FIELD (3 sentences required):\n1. WHAT: what is this record, key evidence from text\n2. WHY: why scores are what they are, cite specific signals\n3. NEXT: what operator should do and why\n\nOUTPUT FORMAT -- CRITICAL:\nRespond with ONLY a valid JSON object.\nNo markdown, no code fences, no preamble.\nFirst character must be {. Last must be }.\nAll 25 fields required. Empty string for unknown strings. 1 for unknown integers.\n\nREQUIRED JSON SCHEMA:\n{\n  \"created_at\": \"<parsed_at value ISO 8601>\",\n  \"source_type\": \"<from input>\",\n  \"platform\": \"<from input>\",\n  \"source_url\": \"<from input>\",\n  \"parsed_at\": \"<from input>\",\n  \"published_at\": \"<from input or empty>\",\n  \"freshness_status\": \"<fresh|recent|old|unknown>\",\n  \"entity_type\": \"<competitor|lead_signal|market_signal|content_idea|irrelevant>\",\n  \"company_name\": \"<explicitly in text only or empty>\",\n  \"profile_name\": \"<explicitly in text only or empty>\",\n  \"profile_url\": \"<from input or empty>\",\n  \"region\": \"<explicitly mentioned or empty>\",\n  \"service_type\": \"<secured_auto_loan|secured_real_estate_loan|pts_loan|refinancing|mortgage_adjacent|generic_lending|unknown>\",\n  \"offer_text\": \"<1 sentence: what offered/sought or content angle title>\",\n  \"terms\": \"<explicit rate/conditions only or empty>\",\n  \"contact_public\": \"<phone/email/Telegram from text only or empty>\",\n  \"text_context\": \"<cleaned summary max 300 chars>\",\n  \"detected_need\": \"<lead_signal only: need+amount+urgency+bank rejection+region or empty>\",\n  \"competitor_strength\": <integer 1-100; 1 if not competitor>,\n  \"lead_signal_score\": <integer 1-100>,\n  \"content_idea_score\": <integer 1-100>,\n  \"quality_score\": <integer 1-100>,\n  \"reason\": \"<3 sentences: what+evidence; why scores; next action>\",\n  \"recommended_action\": \"<monitor|contact|create_content|ignore|investigate>\",\n  \"status\": \"<analyzed|skipped>\"\n}\n\nREMINDER: Return JSON only. No markdown. No analysis outside JSON. For competitor website records, classify entity_type=competitor if the text offers secured lending services, rates, speed, contact, or Moscow/MO coverage.`;\n\nreturn [{ json: {\n  model: 'claude-sonnet-4-6',\n  max_tokens: 1400,\n  temperature: 0.2,\n  system: systemPrompt,\n  messages: [{\n    role: 'user',\n    content: JSON.stringify({\n      source_type: record.source_type || '',\n      platform: record.platform || '',\n      source_url: record.source_url || '',\n      profile_url: record.profile_url || '',\n      parsed_at: record.parsed_at || '',\n      published_at: record.published_at || '',\n      text_context: record.text_context || ''\n    })\n  }]\n}}];"
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000007",
      "name": "Claude Primary API Request",
      "type": "n8n-nodes-base.httpRequest",
      "typeVersion": 4.2,
      "position": [
        2640,
        140
      ],
      "onError": "continueRegularOutput",
      "parameters": {
        "method": "POST",
        "url": "https://aiprimetech.io/v1/messages",
        "authentication": "predefinedCredentialType",
        "nodeCredentialType": "httpHeaderAuth",
        "sendHeaders": true,
        "headerParameters": {
          "parameters": [
            {
              "name": "anthropic-version",
              "value": "2023-06-01"
            }
          ]
        },
        "sendBody": true,
        "contentType": "raw",
        "rawContentType": "application/json",
        "body": "={{ JSON.stringify({ model: $json.model, max_tokens: $json.max_tokens, temperature: $json.temperature, system: $json.system, messages: $json.messages }) }}",
        "options": {}
      },
      "credentials": {
        "httpHeaderAuth": {
          "name": "<your credential>"
        }
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000008",
      "name": "Parse Primary JSON",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        2860,
        140
      ],
      "parameters": {
        "jsCode": "const response = $json;\nconst srcRecord = $('Normalize Firecrawl Output').first().json;\nconst __sd=$getWorkflowStaticData('global');const __r=(__sd.wf04_run=__sd.wf04_run||{});__r.primary_calls=(__r.primary_calls||0)+1;__r.firecrawl_calls=(__r.firecrawl_calls||0)+1;__r.urls_scraped=(__r.urls_scraped||0)+1;__r.claude_calls=(__r.claude_calls||0)+1;function __pp(ok){if(ok)__r.primary_parse_success=(__r.primary_parse_success||0)+1;else __r.primary_parse_failure=(__r.primary_parse_failure||0)+1;}\nfunction cap(s, n) { return (s == null ? '' : String(s)).substring(0, n); }\nfunction extractBalanced(s) {\n  const start = s.indexOf('{');\n  if (start === -1) return null;\n  let depth = 0, inStr = false, esc = false;\n  for (let i = start; i < s.length; i++) {\n    const c = s[i];\n    if (inStr) { if (esc) { esc = false; } else if (c === '\\\\') { esc = true; } else if (c === '\"') { inStr = false; } continue; }\n    if (c === '\"') { inStr = true; continue; }\n    if (c === '{') depth++;\n    else if (c === '}') { depth--; if (depth === 0) return s.substring(start, i + 1); }\n  }\n  return null;\n}\nfunction tryParse(text) {\n  let txt = String(text).trim();\n  txt = txt.replace(/^```json\\s*/i, '').replace(/```\\s*$/, '').trim();\n  txt = txt.replace(/^```\\s*/, '').replace(/```\\s*$/, '').trim();\n  const norm = (x) => x.replace(/[\u2018\u2019]/g, \"'\").replace(/[\u201c\u201d\u00ab\u00bb]/g, '\"');\n  try { return { ok: true, obj: JSON.parse(norm(txt)), candidate: cap(txt, 500) }; } catch(e) {}\n  const bal = extractBalanced(txt);\n  if (bal) { try { return { ok: true, obj: JSON.parse(norm(bal)), candidate: cap(bal, 500) }; } catch(e) {} }\n  const s = txt.indexOf('{'), e2 = txt.lastIndexOf('}');

Credentials you'll need

Each integration node will prompt for credentials when you import. We strip credential IDs before publishing — you'll add your own.

Pro

For the full experience including quality scoring and batch install features for each workflow upgrade to Pro

About this workflow

04 - Firecrawl URL List Mini-Batch to Resilient Analyzer. Uses googleSheets, httpRequest, executeWorkflowTrigger. Event-driven trigger; 42 nodes.

Source: https://github.com/CodeVinci8/vinci-ai-pilot/blob/main/n8n/workflows/04_firecrawl_url_list_resilient.json — original creator credit. Request a take-down →

More Web Scraping workflows → · Browse all categories →

Related workflows

Workflows that share integrations, category, or trigger type with this one. All free to copy and import.

Web Scraping

Automate LinkedIn lead generation by scraping comments from targeted posts and enriching profiles with detailed data

Form Trigger, HTTP Request, Google Sheets
Web Scraping

This automated n8n workflow scrapes job listings from Upwork using Apify, processes and cleans the data, and generates daily email reports with job summaries. The system uses Google Sheets for data st

Google Sheets, HTTP Request, Gmail
Web Scraping

Transform LinkedIn profile URLs into comprehensive enriched lead profiles, quickly and automatically.

HTTP Request, Google Sheets
Web Scraping

Transform any website into a structured knowledge repository with this intelligent crawler that extracts hyperlinks from the homepage, intelligently filters images and content pages, and aggregates fu

HTTP Request, Google Sheets
Web Scraping

Content creators, researchers, educators, and digital marketers who need to discover high-quality YouTube training videos on specific topics. Perfect for building curated learning resource lists, comp

HTTP Request, Google Sheets