{
  "name": "04 - Firecrawl URL List Mini-Batch to Resilient Analyzer",
  "nodes": [
    {
      "id": "fc000000-0000-0000-0000-0000000000n1",
      "name": "Overview Note RU",
      "type": "n8n-nodes-base.stickyNote",
      "typeVersion": 1,
      "position": [
        -360,
        -160
      ],
      "parameters": {
        "content": "## 04 \u2014 Firecrawl: \u0441\u043f\u0438\u0441\u043e\u043a URL (\u043c\u0438\u043d\u0438-\u0431\u0430\u0442\u0447) \u2192 \u0443\u0441\u0442\u043e\u0439\u0447\u0438\u0432\u044b\u0439 \u0430\u043d\u0430\u043b\u0438\u0437\u0430\u0442\u043e\u0440\n\n\u041e\u0431\u0440\u0430\u0431\u0430\u0442\u044b\u0432\u0430\u0435\u0442 \u0420\u0423\u0427\u041d\u041e\u0419 \u0441\u043f\u0438\u0441\u043e\u043a \u0438\u0437 3\u20135 URL \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u043e\u0432 \u043f\u043e \u043e\u0434\u043d\u043e\u043c\u0443.\n\u041d\u0415 \u0440\u0430\u0441\u043f\u0438\u0441\u0430\u043d\u0438\u0435, \u041d\u0415 crawl, \u041d\u0415 batch-\u044d\u043d\u0434\u043f\u043e\u0438\u043d\u0442, \u041d\u0415 \u0431\u043e\u043b\u044c\u0448\u043e\u0439 \u043f\u0430\u0440\u0441\u0438\u043d\u0433.\n\n\u0427\u0442\u043e \u0434\u0435\u043b\u0430\u0435\u0442 \u043d\u0430 \u043a\u0430\u0436\u0434\u044b\u0439 URL:\n1. \u041d\u043e\u0440\u043c\u0430\u043b\u0438\u0437\u0443\u0435\u0442 URL \u0438 \u043f\u0440\u043e\u0432\u0435\u0440\u044f\u0435\u0442 \u0434\u0443\u0431\u043b\u0438\u043a\u0430\u0442 \u0432 \u043e\u0442\u0434\u0435\u043b\u044c\u043d\u043e\u0439 \u0432\u043a\u043b\u0430\u0434\u043a\u0435 url_registry (\u043f\u043e normalized_source_url) \u0414\u041e \u043b\u044e\u0431\u044b\u0445 \u0442\u0440\u0430\u0442.\n2. \u0414\u0443\u0431\u043b\u0438\u043a\u0430\u0442 (\u0438 force_reprocess=false) \u2192 skipped_log, parse_method=dedup_source_url, \u0411\u0415\u0417 Firecrawl/Claude (0 \u0441\u0442\u043e\u0438\u043c\u043e\u0441\u0442\u0438).\n3. \u041d\u043e\u0432\u044b\u0439 URL \u2192 Firecrawl (POST /v2/scrape, markdown) \u2192 \u043e\u0447\u0438\u0441\u0442\u043a\u0430 (\u0431\u0435\u0437 \u043a\u0430\u0440\u0442\u0438\u043d\u043e\u043a/svg, \u0440\u0435\u043b\u0435\u0432\u0430\u043d\u0442\u043d\u043e\u0435 \u043f\u0435\u0440\u0432\u044b\u043c) \u2192 \u0437\u0430\u043f\u0438\u0441\u044c-\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a (text_context \u2264 3500).\n4. \u2192 \u0442\u043e\u0442 \u0436\u0435 \u0443\u0441\u0442\u043e\u0439\u0447\u0438\u0432\u044b\u0439 \u0430\u043d\u0430\u043b\u0438\u0437\u0430\u0442\u043e\u0440 (Claude \u2192 \u0440\u0430\u0437\u0431\u043e\u0440 \u2192 \u0440\u0435\u043c\u043e\u043d\u0442 \u043f\u0440\u0438 \u0441\u0431\u043e\u0435 \u2192 \u0434\u0435\u0442\u0435\u0440\u043c\u0438\u043d\u0438\u0440\u043e\u0432\u0430\u043d\u043d\u044b\u0439 fallback \u043f\u043e \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u0443 \u2192 \u043d\u043e\u0440\u043c\u0430\u043b\u0438\u0437\u0430\u0446\u0438\u044f \u2192 \u043c\u0430\u0440\u0448\u0440\u0443\u0442).\n5. \u2192 \u043d\u0443\u0436\u043d\u0430\u044f \u0432\u043a\u043b\u0430\u0434\u043a\u0430 Google Sheets (Sheet Name = route).\n6. \u041f\u043e\u0441\u043b\u0435 \u043e\u0431\u0440\u0430\u0431\u043e\u0442\u043a\u0438 (\u043d\u0435 \u0434\u0443\u0431\u043b\u0438\u043a\u0430\u0442) \u2192 \u0441\u0442\u0440\u043e\u043a\u0430 \u0432 url_registry (10 \u043a\u043e\u043b\u043e\u043d\u043e\u043a).\n\nFirecrawl \u0443\u043f\u0430\u043b/\u043f\u0443\u0441\u0442\u043e \u2192 technical_errors \u0411\u0415\u0417 Claude (\u043d\u043e \u0432\u0441\u0451 \u0440\u0430\u0432\u043d\u043e \u043f\u0438\u0448\u0435\u0442\u0441\u044f \u0432 url_registry). \u0421\u0431\u043e\u0439 \u043e\u0434\u043d\u043e\u0433\u043e URL \u043d\u0435 \u043e\u0441\u0442\u0430\u043d\u0430\u0432\u043b\u0438\u0432\u0430\u0435\u0442 \u043e\u0441\u0442\u0430\u043b\u044c\u043d\u044b\u0435.\n\u041f\u0440\u0430\u0439\u043c\u0435\u0440\u0438+\u0440\u0435\u043c\u043e\u043d\u0442 \u043d\u0435 \u0434\u0430\u043b\u0438 JSON, \u043d\u043e \u22655 \u0441\u0438\u0433\u043d\u0430\u043b\u043e\u0432 \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u0430 \u2192 deterministic_competitor_fallback \u2192 monitor_queue.\n\n\u041b\u0438\u043c\u0438\u0442\u044b: \u043c\u0430\u043a\u0441\u0438\u043c\u0443\u043c 5 URL (\u0436\u0451\u0441\u0442\u043a\u0430\u044f \u043e\u0442\u0441\u0435\u0447\u043a\u0430), \u043f\u0435\u0440\u0432\u044b\u0439 \u0437\u0430\u043f\u0443\u0441\u043a \u2014 3 URL. \u041d\u0435 \u0430\u043a\u0442\u0438\u0432\u0438\u0440\u043e\u0432\u0430\u0442\u044c. \u0421\u0445\u0435\u043c\u0430 35 \u043a\u043e\u043b\u043e\u043d\u043e\u043a (\u0441 run_id, batch_index); url_registry \u2014 \u0441\u0432\u043e\u0438 10 \u043a\u043e\u043b\u043e\u043d\u043e\u043a.\n\u041e\u0436\u0438\u0434\u0430\u0435\u043c\u043e: \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0430 \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u0430 \u2192 monitor_queue; \u0434\u0443\u0431\u043b\u0438\u043a\u0430\u0442 \u2192 skipped_log.",
        "height": 420,
        "width": 540,
        "color": 4
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000002",
      "name": "Manual Start",
      "type": "n8n-nodes-base.manualTrigger",
      "typeVersion": 1,
      "position": [
        120,
        300
      ],
      "parameters": {}
    },
    {
      "id": "b4000000-0000-0000-0000-00000000000l",
      "name": "Set URL List",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        220,
        300
      ],
      "parameters": {
        "jsCode": "// OPERATOR: edit rawUrls below. MAX 5 URLs (hard cap). First run: use 3.\n// Keep ONLY example.com placeholders committed to the repo \u2014 paste real URLs locally.\nvar __agentIn = (typeof $json === 'object' && $json) ? $json : {};\nconst __exampleUrls = [\n  'https://example.com/competitor-1',\n  'https://example.com/competitor-2',\n  'https://example.com/competitor-3'\n];\nconst rawUrls = (Array.isArray(__agentIn.urls) && __agentIn.urls.length) ? __agentIn.urls : __exampleUrls;\n\nfunction pad(n){ return String(n).padStart(2, '0'); }\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nfunction moscowStamp(){var z=function(n){return String(n).padStart(2,'0');};var m=new Date(Date.now()+10800000);return m.getUTCFullYear()+z(m.getUTCMonth()+1)+z(m.getUTCDate())+'_'+z(m.getUTCHours())+z(m.getUTCMinutes())+z(m.getUTCSeconds());}\nconst now = moscowIsoNow();\nconst __stamp = moscowStamp();\nconst run_id = (__agentIn.source_run_id || __agentIn.run_id) ? String(__agentIn.source_run_id || __agentIn.run_id) : ('firecrawl_' + __stamp);\nconst agent_request_id = __agentIn.agent_request_id ? String(__agentIn.agent_request_id) : ('wf04_req_' + __stamp); // distinct request id (not the source_run_id)\n// SOURCE-REUSE-001: the caller's owner + execution-mode context. WF04 must make the SAME reuse/collect/refresh\n// decision the planner promised, so the planner's inputs travel with every item.\nconst owner_user_id = __agentIn.owner_user_id ? String(__agentIn.owner_user_id) : '';\nconst source_execution_mode = __agentIn.source_execution_mode ? String(__agentIn.source_execution_mode) : '';\nconst freshness_days = Number(__agentIn.freshness_days) || 0;\n\nconst cleaned = [];\nfor (let u of rawUrls) {\n  if (u == null) continue;\n  u = String(u).trim();\n  if (u === '') continue;\n  cleaned.push(u);\n}\nconst capped = cleaned.slice(0, 5); // HARD CAP: max 5 URLs\n\n// --- Stage C Closure Patch 2: reset run-level repair/fallback accounting (S2-D7/D8/D9/D16) ---\nconst __sd=$getWorkflowStaticData('global');\n__sd.wf04_run={ urls_received:capped.length, urls_scraped:0, primary_calls:0, primary_parse_success:0, primary_parse_failure:0, repair_calls:0, repair_success:0, repair_failure:0, deterministic_fallback:0, degraded:0, quarantined:0, snapshots_written:0, firecrawl_calls:0, claude_calls:0, reused:0, reuse_failed:0, original_snapshot_run_id:'', original_snapshot_collected_at:'', force_reprocess:(__agentIn.force_reprocess === true || __agentIn.force_reprocess === 'true'), requested_mode:source_execution_mode };\nconst items = [];\nlet idx = 0;\nfor (const u of capped) {\n  idx++;\n  items.push({ json: {\n    target_url: u,\n    source_type: 'scraped_web',\n    platform: 'website',\n    parsed_at: now,\n    source_note: 'firecrawl_url_list_manual',\n    run_id: run_id,\n    agent_request_id: agent_request_id,\n    batch_index: idx,\n    force_reprocess: (__agentIn.force_reprocess === true || __agentIn.force_reprocess === 'true'), // FORCE-REPROCESS-001: callable passes strings\n    owner_user_id: owner_user_id,\n    source_execution_mode: source_execution_mode,\n    freshness_days: freshness_days\n  }});\n}\nreturn items;"
      }
    },
    {
      "id": "b4000000-0000-0000-0000-00000000000b",
      "name": "Loop Over Items",
      "type": "n8n-nodes-base.splitInBatches",
      "typeVersion": 3,
      "position": [
        440,
        300
      ],
      "parameters": {
        "options": {}
      }
    },
    {
      "id": "b4000000-0000-0000-0000-00000000000n",
      "name": "Normalize URL for Dedup",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        660,
        300
      ],
      "parameters": {
        "jsCode": "const j = $json;\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nlet u = String(j.target_url || '').trim();\nlet normalized = u;\ntry {\n  if (u) {\n    const noFrag = u.split('#')[0];\n    const url = new URL(noFrag);\n    url.protocol = (url.protocol || '').toLowerCase();\n    url.hostname = (url.hostname || '').toLowerCase();\n    const drop = ['utm_source','utm_medium','utm_campaign','utm_term','utm_content','gclid','yclid','fbclid'];\n    for (const p of drop) url.searchParams.delete(p);\n    let path = url.pathname || '/';\n    if (path.length > 1 && path.endsWith('/')) path = path.slice(0, -1);\n    url.pathname = path;\n    normalized = url.toString();\n  }\n} catch(e) {\n  normalized = u;\n}\nreturn [{ json: {\n  target_url: u || normalized,\n  normalized_source_url: normalized,\n  source_url: normalized,\n  source_type: j.source_type || 'scraped_web',\n  platform: j.platform || 'website',\n  parsed_at: j.parsed_at || moscowIsoNow(),\n  source_note: j.source_note || 'firecrawl_url_list_manual',\n  run_id: j.run_id || '',\n  batch_index: j.batch_index || 0,\n  force_reprocess: (j.force_reprocess === true || j.force_reprocess === 'true'),\n  owner_user_id: j.owner_user_id || '',\n  source_execution_mode: j.source_execution_mode || '',\n  freshness_days: Number(j.freshness_days) || 0\n}}];"
      }
    },
    {
      "id": "b4000000-0000-0000-0000-0000000000rg",
      "name": "Registry Lookup",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        880,
        300
      ],
      "alwaysOutputData": true,
      "onError": "continueRegularOutput",
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "read",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "url_registry",
          "mode": "name"
        },
        "filtersUI": {
          "values": [
            {
              "lookupColumn": "normalized_source_url",
              "lookupValue": "={{ $('Normalize URL for Dedup').first().json.normalized_source_url }}"
            }
          ]
        },
        "options": {
          "returnAllMatches": "returnAllMatches"
        }
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "b4000000-0000-0000-0000-00000000000e",
      "name": "Evaluate Dedup",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1100,
        300
      ],
      "parameters": {
        "jsCode": "const ctx = $('Normalize URL for Dedup').first().json;\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nconst now = moscowIsoNow();\nconst key = String(ctx.normalized_source_url || '').trim();\nconst run_id = ctx.run_id || '';\nconst batch_index = ctx.batch_index || 0;\nconst parsedAt = ctx.parsed_at || now;\nconst force = (ctx.force_reprocess === true || ctx.force_reprocess === 'true');\nconst ownerId = String(ctx.owner_user_id || '');\nconst freshnessDays = Number(ctx.freshness_days) || 0;\n\n// embedded n8n/lib/source_execution_policy.js (do not edit here; edit the lib and re-run the transform)\n// source_execution_policy.js \u2014 the ONE decision for \"do we pay to collect this source again?\"\n//\n// SOURCE-EXEC-001. Before this, WF04's `url_registry` check was a PERMANENT dedup: `hit = !force && rows.some(...)`\n// with no time component. Once a URL was ever scraped it was skipped forever, so a user could never re-analyze a\n// site \u2014 and the run produced an empty bundle and a misleading \"\u0434\u0430\u043d\u043d\u044b\u0445 \u043d\u0435\u0442\" while a perfectly good saved snapshot\n// sat in the sheet. Permanent-skip is not a freshness policy; it is a leak.\n//\n// Three explicit modes:\n//   reuse   \u2014 a recent ACCEPTED snapshot exists and is still fresh. No paid collection; analyze the saved snapshot\n//             and say so, with its collection time. Collection cost $0.\n//   collect \u2014 no accepted snapshot, or the newest is older than the TTL, or every prior attempt failed. Pay once.\n//   refresh \u2014 the user explicitly asked to re-collect. Bypasses ONLY the freshness check; approval, budget, quality\n//             and global dedup all still apply, and the repeated paid collection is stated in the plan.\n//\n// A FAILED snapshot never counts as reusable and never blocks a retry \u2014 otherwise one bad scrape would poison a\n// source forever (the same permanence bug in a different costume).\n//\n// Embeddable: unique sx*-prefixed names, no cross-lib require.\n\nfunction sxStr(v) { return v == null ? '' : String(v); }\nfunction sxNum(v, d) { var n = Number(v); return isFinite(n) ? n : (d === undefined ? 0 : d); }\n\nvar SX_MODES = { REUSE: 'reuse', COLLECT: 'collect', REFRESH: 'refresh' };\n// Default freshness window. The product's monitoring/report cadence is weekly (WF25 weekly digest, WF23 monitor),\n// and WF10/WF12 already reason over a 30-day window \u2014 so a snapshot younger than 7 days is \"current\" for a\n// competitor's public positioning (offers/prices change on a weeks-to-months scale, not hourly). Operator override:\n// MS_SOURCE_FRESHNESS_DAYS.\nvar SX_DEFAULT_FRESHNESS_DAYS = 7;\n\n// A snapshot is reusable only if it actually produced accepted content.\nvar SX_REJECTED_STATUSES = ['failed', 'error', 'quarantined', 'technical_error', 'excluded', 'invalid'];\nfunction sxIsAccepted(s) {\n  if (!s) return false;\n  var st = sxStr(s.quality_status || s.status).toLowerCase();\n  if (st && SX_REJECTED_STATUSES.indexOf(st) >= 0) return false;\n  // SOURCE-REUSE-001: url_registry rows carry the verdict in processing_status (last_processing_status), not in\n  // quality_status. Checking only technical_error here let a quarantined/failed prior run count as \"accepted\".\n  var ps = sxStr(s.processing_status).toLowerCase();\n  if (ps && SX_REJECTED_STATUSES.indexOf(ps) >= 0) return false;\n  if (s.accepted === false) return false;\n  return true;\n}\nfunction sxTime(s) {\n  var t = Date.parse(sxStr((s && (s.collected_at || s.parsed_at || s.last_seen_at || s.created_at)) || ''));\n  return isFinite(t) ? t : NaN;\n}\nfunction sxNorm(u) {\n  return sxStr(u).trim().toLowerCase().replace(/^https?:\\/\\//, '').replace(/^www\\./, '').replace(/\\/+$/, '');\n}\n\n// newestAcceptedSnapshot(snapshots, source_url, owner) -> the freshest ACCEPTED snapshot for this url (or null).\n// Owner isolation: when a snapshot carries an owner, only that owner's rows are considered.\nfunction newestAcceptedSnapshot(snapshots, sourceUrl, owner) {\n  var key = sxNorm(sourceUrl);\n  var best = null, bestT = -1;\n  (Array.isArray(snapshots) ? snapshots : []).forEach(function (s) {\n    if (!s) return;\n    if (sxNorm(s.source_url || s.normalized_source_url) !== key) return;\n    if (owner && sxStr(s.owner_user_id) && sxStr(s.owner_user_id) !== sxStr(owner)) return;\n    if (!sxIsAccepted(s)) return;\n    var t = sxTime(s);\n    if (isNaN(t)) return;\n    if (t > bestT) { bestT = t; best = s; }\n  });\n  return best;\n}\n\n// decideSourceExecution(input) -> { mode, reason, force_reprocess, snapshot, snapshot_age_days, snapshot_collected_at }\n// input: { source_url, snapshots[], owner_user_id, requested_refresh, now, freshness_days }\nfunction decideSourceExecution(input) {\n  input = input || {};\n  var nowMs = input.now ? Date.parse(sxStr(input.now)) : Date.now();\n  if (!isFinite(nowMs)) nowMs = Date.now();\n  var ttlDays = sxNum(input.freshness_days, SX_DEFAULT_FRESHNESS_DAYS);\n  if (!(ttlDays > 0)) ttlDays = SX_DEFAULT_FRESHNESS_DAYS;\n\n  var snap = newestAcceptedSnapshot(input.snapshots, input.source_url, input.owner_user_id);\n  // Did we try this url before and fail? That is NOT \"never collected\" \u2014 the user deserves the real reason, and a\n  // failed attempt must never block a retry.\n  var key = sxNorm(input.source_url);\n  var triedBefore = (Array.isArray(input.snapshots) ? input.snapshots : []).some(function (s) {\n    return s && sxNorm(s.source_url || s.normalized_source_url) === key &&\n      (!input.owner_user_id || !sxStr(s.owner_user_id) || sxStr(s.owner_user_id) === sxStr(input.owner_user_id));\n  });\n  var ageDays = snap ? (nowMs - sxTime(snap)) / 86400000 : null;\n  var out = {\n    mode: SX_MODES.COLLECT, reason: triedBefore ? 'last_attempt_failed' : 'never_collected', force_reprocess: false,\n    snapshot: null, snapshot_age_days: null, snapshot_collected_at: ''\n  };\n  if (snap) {\n    out.snapshot = snap;\n    out.snapshot_age_days = Math.round(ageDays * 100) / 100;\n    out.snapshot_collected_at = sxStr(snap.collected_at || snap.parsed_at || snap.last_seen_at || snap.created_at);\n  }\n\n  // An explicit refresh wins over freshness \u2014 but ONLY over freshness. Everything else still gates it.\n  if (input.requested_refresh === true) {\n    out.mode = SX_MODES.REFRESH;\n    out.force_reprocess = true;\n    out.reason = snap ? 'explicit_refresh' : 'explicit_refresh_no_snapshot';\n    return out;\n  }\n  if (!snap) return out;                                   // collect / never_collected\n  if (ageDays > ttlDays) { out.reason = 'snapshot_stale'; return out; }  // collect / stale\n  out.mode = SX_MODES.REUSE;\n  out.reason = 'fresh_snapshot';\n  return out;\n}\n\n// Russian phrases that mean \"collect it again, now\". Cyrillic \\b/\\w do not fire in JS, so match on explicit\n// [\u0430-\u044f\u0451] boundaries. Deliberately narrow: an accidental refresh costs the user real money.\nvar SX_REFRESH_RE = /(^|[^\u0430-\u044f\u0451a-z])(\u043e\u0431\u043d\u043e\u0432\u0438(\u0442\u044c|\u0442\u0435)?|\u043f\u0435\u0440\u0435\u043e\u0431\u043d\u043e\u0432\u0438(\u0442\u044c|\u0442\u0435)?|\u043f\u0435\u0440\u0435\u0441\u043e\u0431\u0435\u0440\u0438|\u043f\u0435\u0440\u0435\u0441\u043e\u0431\u0440\u0430\u0442\u044c|\u043f\u0435\u0440\u0435\u0441\u043e\u0431\u0435\u0440\u0438\u0442\u0435|\u0437\u0430\u043d\u043e\u0432\u043e|\u043f\u043e\u0432\u0442\u043e\u0440\u0438(\u0442\u044c|\u0442\u0435)? \u0441\u0431\u043e\u0440|\u043f\u043e\u0432\u0442\u043e\u0440\u043d\u044b\u0439 \u0441\u0431\u043e\u0440|\u043f\u0440\u0438\u043d\u0443\u0434\u0438\u0442\u0435\u043b\u044c\u043d(\u043e|\u044b\u0439|\u0430\u044f)|\u0435\u0449\u0451 \u0440\u0430\u0437 \u0441\u043e\u0431\u0435\u0440\u0438|\u0435\u0449\u0435 \u0440\u0430\u0437 \u0441\u043e\u0431\u0435\u0440\u0438|\u0441\u0432\u0435\u0436(\u0438\u0435|\u0438\u0445) \u0434\u0430\u043d\u043d(\u044b\u0435|\u044b\u0445)|\u0430\u043a\u0442\u0443\u0430\u043b\u0438\u0437\u0438\u0440\u0443\u0439|\u043f\u0435\u0440\u0435\u043f\u0440\u043e\u0432\u0435\u0440\u044c)([^\u0430-\u044f\u0451a-z]|$)/i;\n// \"\u043e\u0431\u043d\u043e\u0432\u0438 \u043e\u0442\u0447\u0451\u0442\" = rebuild the report from what we have; it is NOT automatically a paid re-collection.\nvar SX_REPORT_ONLY_RE = /\u043e\u0431\u043d\u043e\u0432\u0438(\u0442\u044c|\u0442\u0435)?\\s+\u043e\u0442\u0447[\u0435\u0451]\u0442/i;\n\n// detectRefreshRequest(text) -> { requested_refresh, refresh_reason }\nfunction detectRefreshRequest(text) {\n  var t = sxStr(text);\n  if (!t) return { requested_refresh: false, refresh_reason: '' };\n  if (SX_REPORT_ONLY_RE.test(t) && !/\u0434\u0430\u043d\u043d|\u0441\u0431\u043e\u0440|\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a|\u0441\u0430\u0439\u0442/i.test(t)) {\n    return { requested_refresh: false, refresh_reason: 'report_rebuild_only' };\n  }\n  if (SX_REFRESH_RE.test(t)) return { requested_refresh: true, refresh_reason: 'user_requested_refresh' };\n  return { requested_refresh: false, refresh_reason: '' };\n}\n// --- end embedded source_execution_policy ---\n\n// SOURCE-REUSE-001. The old check here was a PERMANENT dedup: any registry hit emitted only a skipped_log row and\n// ZERO data, so the planner could promise \"reuse\" while the executor delivered \"no_sources\" (live: WF04 exec 948 ->\n// WF10 exec 951 rows_after_isolation=0 for a site scraped 25 minutes earlier). Planning and execution now share the\n// ONE canonical decision: decideSourceExecution (reuse | collect | refresh).\nlet rows = [];\ntry { rows = $('Registry Lookup').all().map(i => i.json); } catch(e) { rows = []; }\n// url_registry rows -> the policy's snapshot shape (the registry's last_processing_status IS the snapshot status).\nconst snapshots = rows.filter(r => r).map(function (r) { return {\n  source_url: r.normalized_source_url || r.source_url,\n  normalized_source_url: r.normalized_source_url,\n  last_seen_at: r.last_seen_at || r.first_seen_at,\n  processing_status: r.last_processing_status,\n  run_id: r.run_id,\n  last_route: r.last_route,\n  owner_user_id: r.owner_user_id // absent on today's registry rows; enforced whenever present\n};});\nconst decision = decideSourceExecution({ source_url: key, snapshots: snapshots, owner_user_id: ownerId, requested_refresh: force, freshness_days: freshnessDays, now: now });\n\n// Reuse is only real if the original run left something WF10 can actually consume: a row in one of the three data\n// queues, from a run whose source_health verdict is still eligible. Anything else -> collect again (fail closed:\n// a blocked/quarantined/unscored original never silently becomes \"saved data\").\nconst REUSABLE_ROUTES = ['monitor_queue', 'content_queue', 'review_queue'];\nlet originalHealth = null;\nif (decision.mode === 'reuse') {\n  const origRoute = String((decision.snapshot && decision.snapshot.last_route) || '');\n  if (REUSABLE_ROUTES.indexOf(origRoute) < 0) {\n    decision.mode = 'collect';\n    decision.reason = 'previous_run_not_reusable_route';\n  } else {\n    let healthRows = [];\n    try { healthRows = $('Source Health Lookup').all().map(i => i.json); } catch(e) { healthRows = []; }\n    const orig = String(decision.snapshot.run_id || '');\n    const BAD_HEALTH = ['quarantined', 'failed', 'error', 'invalid', 'excluded'];\n    healthRows.forEach(function (h) {\n      if (!h || String(h.source_run_id || '') !== orig) return;\n      const t = Date.parse(String(h.evaluated_at || '')) || 0;\n      if (!originalHealth || t >= (Date.parse(String(originalHealth.evaluated_at || '')) || 0)) originalHealth = h;\n    });\n    const hOk = originalHealth\n      && BAD_HEALTH.indexOf(String(originalHealth.quality_status || '').toLowerCase()) < 0\n      && String(originalHealth.report_eligible).toLowerCase() !== 'false';\n    if (!hOk) {\n      decision.mode = 'collect';\n      decision.reason = originalHealth ? 'original_run_not_eligible' : 'original_run_unscored';\n      originalHealth = null;\n    }\n  }\n}\n\nif (decision.mode === SX_MODES.REUSE) {\n  return [{ json: {\n    sx_mode: 'reuse',\n    sx_reason: decision.reason,\n    target_url: ctx.target_url || key,\n    normalized_source_url: key,\n    source_url: key,\n    source_type: ctx.source_type || 'scraped_web',\n    platform: ctx.platform || 'website',\n    parsed_at: parsedAt,\n    run_id: run_id,\n    batch_index: batch_index,\n    owner_user_id: ownerId,\n    reuse_route: String(decision.snapshot.last_route),\n    original_run_id: String(decision.snapshot.run_id || ''),\n    original_collected_at: String(decision.snapshot_collected_at || ''),\n    original_snapshot_age_days: decision.snapshot_age_days,\n    original_health_row: originalHealth\n  }}];\n}\n\n// collect / refresh -> Firecrawl branch (same record shape as before, now with the typed decision attached).\nreturn [{ json: {\n  sx_mode: decision.mode,\n  sx_reason: decision.reason,\n  target_url: ctx.target_url || key,\n  normalized_source_url: key,\n  source_url: key,\n  source_type: ctx.source_type || 'scraped_web',\n  platform: ctx.platform || 'website',\n  parsed_at: parsedAt,\n  run_id: run_id,\n  batch_index: batch_index\n}}];"
      }
    },
    {
      "id": "b4000000-0000-0000-0000-00000000000i",
      "name": "IF Reuse?",
      "type": "n8n-nodes-base.if",
      "typeVersion": 2,
      "position": [
        1320,
        300
      ],
      "parameters": {
        "conditions": {
          "options": {
            "caseSensitive": true,
            "leftValue": "",
            "typeValidation": "loose"
          },
          "conditions": [
            {
              "id": "b4-if-dup-01",
              "leftValue": "={{ $json.sx_mode }}",
              "rightValue": "reuse",
              "operator": {
                "type": "string",
                "operation": "equals"
              }
            }
          ],
          "combinator": "and"
        }
      }
    },
    {
      "id": "fc000000-0000-0000-0000-000000000002",
      "name": "Build Firecrawl Request",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1540,
        300
      ],
      "parameters": {
        "jsCode": "const url = $json.target_url || '';\nreturn [{ json: {\n  url: url,\n  formats: ['markdown'],\n  onlyMainContent: true,\n  onlyCleanContent: false,\n  removeBase64Images: true,\n  blockAds: true,\n  timeout: 60000,\n  storeInCache: true\n}}];"
      }
    },
    {
      "id": "fc000000-0000-0000-0000-000000000003",
      "name": "Firecrawl Scrape API",
      "type": "n8n-nodes-base.httpRequest",
      "typeVersion": 4.2,
      "position": [
        1760,
        300
      ],
      "onError": "continueRegularOutput",
      "parameters": {
        "method": "POST",
        "url": "https://api.firecrawl.dev/v2/scrape",
        "authentication": "predefinedCredentialType",
        "nodeCredentialType": "httpHeaderAuth",
        "sendHeaders": true,
        "headerParameters": {
          "parameters": [
            {
              "name": "Content-Type",
              "value": "application/json"
            }
          ]
        },
        "sendBody": true,
        "contentType": "raw",
        "rawContentType": "application/json",
        "body": "={{ JSON.stringify($json) }}",
        "options": {}
      },
      "credentials": {
        "httpHeaderAuth": {
          "name": "<your credential>"
        }
      }
    },
    {
      "id": "fc000000-0000-0000-0000-000000000004",
      "name": "Normalize Firecrawl Output",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1980,
        300
      ],
      "parameters": {
        "jsCode": "// embedded n8n/lib/source_access.js (drift-proof; test asserts equality)\n// source_access.js \u2014 BLOCK-HONESTY-001. Decide whether we actually REACHED a source, BEFORE anyone asks whether the\n// business is relevant.\n//\n// The bug this exists to kill (live: carmoney.ru, WF04 exec 894): Firecrawl returned a 772-char access-restriction\n// page. It was longer than the 80-char \"meaningful content\" floor, so it flowed into business classification, which\n// \u2014 correctly, given the text it saw \u2014 said `entity_type=irrelevant / business_skip`. The user was then told\n// \u00abcarmoney.ru \u2014 \u043f\u0440\u043e\u0432\u0435\u0440\u0435\u043d, \u043d\u043e\u0432\u044b\u0445 \u0440\u0435\u043b\u0435\u0432\u0430\u043d\u0442\u043d\u044b\u0445 \u0444\u0430\u043a\u0442\u043e\u0432 \u043d\u0435 \u043d\u0430\u0439\u0434\u0435\u043d\u043e\u00bb and advised to widen filters. Every part of that is\n// false: we never saw the site, the company may be highly relevant, and no filter change can fix an IP block.\n//\n// \"We could not read the page\" and \"we read the page and it is not a competitor\" are DIFFERENT facts with different\n// user messages, different next actions, and different data consequences. Access is decided first, from the\n// transport + the page's own text; relevance is only asked once access is `accessible_content`.\n//\n// A non-accessible outcome may NEVER become a competitor snapshot, overwrite a good snapshot, or feed Claude.\n//\n// Embeddable: unique sa*-prefixed names, no cross-lib require.\n\nfunction saStr(v) { return v == null ? '' : String(v); }\nfunction saNum(v, d) { var n = Number(v); return isFinite(n) ? n : (d === undefined ? 0 : d); }\n\nvar SA_OUTCOMES = {\n  ACCESSIBLE: 'accessible_content',\n  BLOCKED_WAF: 'blocked_by_waf',\n  ACCESS_DENIED: 'robots_or_access_denied',\n  PROVIDER_FAILURE: 'provider_failure',\n  TIMEOUT: 'timeout',\n  EMPTY: 'empty_response',\n  UNSUPPORTED: 'unsupported_content',\n  IRRELEVANT: 'valid_but_irrelevant'   // set DOWNSTREAM, only after ACCESSIBLE \u2014 never inferred here\n};\n// Only these mean \"we hold real page content\".\nfunction saIsAccessible(o) { return o === SA_OUTCOMES.ACCESSIBLE; }\n// These are transport/access problems: the business is unjudged, so the user must never be told it is irrelevant.\nfunction saIsAccessFailure(o) {\n  return [SA_OUTCOMES.BLOCKED_WAF, SA_OUTCOMES.ACCESS_DENIED, SA_OUTCOMES.PROVIDER_FAILURE,\n    SA_OUTCOMES.TIMEOUT, SA_OUTCOMES.EMPTY, SA_OUTCOMES.UNSUPPORTED].indexOf(o) >= 0;\n}\n// Worth another attempt later (a block/timeout may lift); an unsupported content type will not fix itself.\nfunction saIsRetryable(o) {\n  return [SA_OUTCOMES.BLOCKED_WAF, SA_OUTCOMES.PROVIDER_FAILURE, SA_OUTCOMES.TIMEOUT, SA_OUTCOMES.EMPTY].indexOf(o) >= 0;\n}\n\n// STRONG signatures: unambiguous challenge/block boilerplate. A real commercial page does not say these about itself.\nvar SA_STRONG_WAF = [\n  'just a moment...', 'checking your browser before accessing', 'attention required! | cloudflare',\n  'cloudflare ray id', 'enable javascript and cookies to continue', 'verify you are human',\n  'ddos protection by cloudflare', 'performance & security by cloudflare', 'error 1020', 'error code 1020',\n  'ray id:', 'sorry, you have been blocked', 'why have i been blocked', 'incapsula incident id',\n  'request unsuccessful. incapsula', 'access to this page has been denied', 'pardon our interruption',\n  'are you a robot', '\u043f\u043e\u0434\u0442\u0432\u0435\u0440\u0434\u0438\u0442\u0435, \u0447\u0442\u043e \u0432\u044b \u043d\u0435 \u0440\u043e\u0431\u043e\u0442', '\u043f\u0440\u043e\u0432\u0435\u0440\u043a\u0430 \u0431\u0440\u0430\u0443\u0437\u0435\u0440\u0430', '\u0434\u043e\u0441\u0442\u0443\u043f \u043e\u0433\u0440\u0430\u043d\u0438\u0447\u0435\u043d'\n];\nvar SA_STRONG_DENIED = [\n  '403 forbidden', 'http error 403', 'access denied', 'you don\\'t have permission to access',\n  'you do not have permission to access', 'authorization required', '401 unauthorized',\n  '\u0434\u043e\u0441\u0442\u0443\u043f \u0437\u0430\u043f\u0440\u0435\u0449\u0451\u043d', '\u0434\u043e\u0441\u0442\u0443\u043f \u0437\u0430\u043f\u0440\u0435\u0449\u0435\u043d', '\u043d\u0435\u0442 \u0434\u043e\u0441\u0442\u0443\u043f\u0430 \u043a \u044d\u0442\u043e\u0439 \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0435'\n];\n// The live carmoney.ru family: an IP/geo restriction notice. Two independent phrases must co-occur, so a page that\n// merely mentions \"VPN\" as a product never trips it.\nvar SA_GEOBLOCK_PAIRS = [\n  ['is not available', 'restricted access from your current network'],\n  ['is not available', 'block access from specific countries'],\n  ['restricted access from your current network', 'enabled a vpn'],\n  ['website owner has restricted access', 'ip addresses'],\n  ['\u043d\u0435\u0434\u043e\u0441\u0442\u0443\u043f\u0435\u043d', '\u043e\u0433\u0440\u0430\u043d\u0438\u0447\u0438\u043b \u0434\u043e\u0441\u0442\u0443\u043f']\n];\n// WEAK signals: only meaningful on a SHORT page with no commercial content of its own.\nvar SA_WEAK = ['captcha', 'recaptcha', 'hcaptcha', 'challenge-platform', 'cf-browser-verification',\n  'security check', 'bot detection', 'unusual traffic', 'rate limited', 'too many requests'];\n\n// A page that talks about lending/pricing/contacts is a real page, whatever boilerplate it also contains.\nvar SA_BUSINESS_TERMS = ['\u043a\u0440\u0435\u0434\u0438\u0442', '\u0437\u0430\u0439\u043c', '\u0437\u0430\u043b\u043e\u0433', '\u043f\u0442\u0441', '\u0441\u0442\u0430\u0432\u043a\u0430', '\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441', '\u0438\u043f\u043e\u0442\u0435\u043a', '\u043e\u0434\u043e\u0431\u0440\u0435\u043d', '\u0437\u0430\u044f\u0432\u043a',\n  '\u0442\u0430\u0440\u0438\u0444', '\u0443\u0441\u043b\u0443\u0433', '\u043e\u0444\u043e\u0440\u043c\u0438\u0442\u044c', '\u0440\u0443\u0431', '\u043f\u0440\u043e\u0446\u0435\u043d\u0442', '\u043e\u0444\u0438\u0441', 'loan', 'credit', 'rate', 'apply'];\n\nfunction saHasBusiness(low) {\n  var n = 0;\n  for (var i = 0; i < SA_BUSINESS_TERMS.length; i++) if (low.indexOf(SA_BUSINESS_TERMS[i]) >= 0) n++;\n  return n;\n}\nfunction saAnyHit(low, list) {\n  for (var i = 0; i < list.length; i++) if (low.indexOf(list[i]) >= 0) return list[i];\n  return '';\n}\nfunction saPairHit(low, pairs) {\n  for (var i = 0; i < pairs.length; i++) {\n    if (low.indexOf(pairs[i][0]) >= 0 && low.indexOf(pairs[i][1]) >= 0) return pairs[i].join(' + ');\n  }\n  return '';\n}\n\n// classifySourceAccess(input) -> { outcome, reason, signature, retryable, access_failure, meaningful_chars, business_terms }\n// input: { status, body_text, error, error_category, content_type, url }\n// Order matters: transport verdicts first (they are authoritative), then the page's own text.\nfunction classifySourceAccess(input) {\n  input = input || {};\n  var status = saNum(input.status, NaN);\n  var text = saStr(input.body_text);\n  var low = text.toLowerCase();\n  var meaningful = text.replace(/[#>*_|[\\]()]/g, ' ').replace(/\\s+/g, ' ').trim();\n  var biz = saHasBusiness(low);\n  var out = function (outcome, reason, signature) {\n    return {\n      outcome: outcome, reason: reason, signature: saStr(signature),\n      retryable: saIsRetryable(outcome), access_failure: saIsAccessFailure(outcome),\n      meaningful_chars: meaningful.length, business_terms: biz\n    };\n  };\n\n  // 1. transport-level truth\n  var ec = saStr(input.error_category).toLowerCase();\n  if (ec === 'timeout' || /timeout|etimedout|timed out/i.test(saStr(input.error))) return out(SA_OUTCOMES.TIMEOUT, 'provider_timeout', ec || 'timeout');\n  if (input.error) return out(SA_OUTCOMES.PROVIDER_FAILURE, 'provider_error', saStr(input.error).slice(0, 60));\n  if (isFinite(status)) {\n    if (status === 401 || status === 403) return out(SA_OUTCOMES.ACCESS_DENIED, 'http_' + status, 'http_' + status);\n    if (status === 429) return out(SA_OUTCOMES.BLOCKED_WAF, 'http_429_rate_limited', 'http_429');\n    if (status >= 500) return out(SA_OUTCOMES.PROVIDER_FAILURE, 'http_' + status, 'http_' + status);\n    if (status >= 400) return out(SA_OUTCOMES.PROVIDER_FAILURE, 'http_' + status, 'http_' + status);\n  }\n  // 2. content type we cannot analyze\n  var ct = saStr(input.content_type).toLowerCase();\n  if (ct && !/text\\/html|text\\/plain|application\\/xhtml|markdown|application\\/json/.test(ct)) {\n    return out(SA_OUTCOMES.UNSUPPORTED, 'unsupported_content_type', ct.slice(0, 40));\n  }\n  // 3. nothing came back\n  if (!meaningful) return out(SA_OUTCOMES.EMPTY, 'empty_body', '');\n\n  // 4. the page's own text says we were blocked. Strong signatures win regardless of length: real commercial pages\n  //    do not describe themselves as a browser challenge.\n  var s = saAnyHit(low, SA_STRONG_WAF);\n  if (s) return out(SA_OUTCOMES.BLOCKED_WAF, 'waf_challenge_page', s);\n  var g = saPairHit(low, SA_GEOBLOCK_PAIRS);\n  if (g) return out(SA_OUTCOMES.BLOCKED_WAF, 'network_or_geo_restriction', g);\n  var d = saAnyHit(low, SA_STRONG_DENIED);\n  if (d) return out(SA_OUTCOMES.ACCESS_DENIED, 'access_denied_page', d);\n\n  // 5. weak signals: only on a short page that carries no commercial content of its own \u2014 otherwise a lender that\n  //    happens to mention \"captcha\" in its FAQ would be wrongly reported as blocked.\n  if (meaningful.length < 1200 && biz === 0) {\n    var w = saAnyHit(low, SA_WEAK);\n    if (w) return out(SA_OUTCOMES.BLOCKED_WAF, 'short_page_block_signature', w);\n  }\n  // 6. too little to analyze (mirrors WF04's existing floor)\n  if (meaningful.length < 80) return out(SA_OUTCOMES.EMPTY, 'body_too_short', String(meaningful.length) + ' chars');\n\n  return out(SA_OUTCOMES.ACCESSIBLE, 'content_ok', '');\n}\n\n// The ONE user-facing Russian sentence per access failure. Never a status code, provider name, or raw page text.\n// Every one states explicitly that the business was NOT judged \u2014 that is the whole point of BLOCK-HONESTY-001.\nvar SA_USER_RU = {\n  blocked_by_waf: '\u0441\u0430\u0439\u0442 \u0432\u0435\u0440\u043d\u0443\u043b \u0437\u0430\u0449\u0438\u0442\u043d\u0443\u044e \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0443 \u0438 \u043d\u0435 \u043e\u0442\u0434\u0430\u043b \u0441\u043e\u0434\u0435\u0440\u0436\u0438\u043c\u043e\u0435 \u2014 \u043f\u0440\u043e\u0447\u0438\u0442\u0430\u0442\u044c \u0435\u0433\u043e \u043d\u0435 \u0443\u0434\u0430\u043b\u043e\u0441\u044c',\n  robots_or_access_denied: '\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a \u0437\u0430\u043a\u0440\u044b\u043b \u0434\u043e\u0441\u0442\u0443\u043f \u043a \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0435',\n  provider_failure: '\u0441\u0435\u0440\u0432\u0438\u0441 \u0441\u0431\u043e\u0440\u0430 \u0434\u0430\u043d\u043d\u044b\u0445 \u043d\u0435 \u0441\u043c\u043e\u0433 \u043f\u043e\u043b\u0443\u0447\u0438\u0442\u044c \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0443',\n  timeout: '\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a \u043d\u0435 \u043e\u0442\u0432\u0435\u0442\u0438\u043b \u0432\u043e\u0432\u0440\u0435\u043c\u044f',\n  empty_response: '\u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0430 \u043e\u0442\u043a\u0440\u044b\u043b\u0430\u0441\u044c \u043f\u0443\u0441\u0442\u043e\u0439 \u2014 \u0441\u043e\u0434\u0435\u0440\u0436\u0438\u043c\u043e\u0433\u043e \u0434\u043b\u044f \u0430\u043d\u0430\u043b\u0438\u0437\u0430 \u043d\u0435\u0442',\n  unsupported_content: '\u043f\u043e \u044d\u0442\u043e\u043c\u0443 \u0430\u0434\u0440\u0435\u0441\u0443 \u043d\u0435 \u0442\u0435\u043a\u0441\u0442\u043e\u0432\u0430\u044f \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0430 \u2014 \u0430\u043d\u0430\u043b\u0438\u0437\u0438\u0440\u043e\u0432\u0430\u0442\u044c \u043d\u0435\u0447\u0435\u0433\u043e'\n};\nfunction saUserMessageRu(outcome) { return SA_USER_RU[saStr(outcome)] || '\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a \u0441\u0435\u0439\u0447\u0430\u0441 \u043d\u0435\u0434\u043e\u0441\u0442\u0443\u043f\u0435\u043d'; }\n\n// Cause-specific next actions (\u00a77). Never \"\u0440\u0430\u0441\u0448\u0438\u0440\u044c\u0442\u0435 \u0444\u0438\u043b\u044c\u0442\u0440\u044b\" for an access failure \u2014 no filter reaches a blocked page.\nvar SA_NEXT_RU = {\n  blocked_by_waf: ['\u043f\u043e\u0432\u0442\u043e\u0440\u0438\u0442\u044c \u043f\u043e\u043f\u044b\u0442\u043a\u0443 \u043f\u043e\u0437\u0436\u0435', '\u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u044c \u043f\u043e\u0441\u043b\u0435\u0434\u043d\u0438\u0439 \u0441\u043e\u0445\u0440\u0430\u043d\u0451\u043d\u043d\u044b\u0439 \u0441\u043d\u0438\u043c\u043e\u043a', '\u043f\u0440\u043e\u0432\u0435\u0440\u0438\u0442\u044c \u0434\u0440\u0443\u0433\u043e\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a'],\n  robots_or_access_denied: ['\u043f\u0440\u043e\u0432\u0435\u0440\u0438\u0442\u044c \u0434\u0440\u0443\u0433\u043e\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a', '\u0438\u0441\u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u044c \u043f\u043e\u0441\u043b\u0435\u0434\u043d\u0438\u0439 \u0441\u043e\u0445\u0440\u0430\u043d\u0451\u043d\u043d\u044b\u0439 \u0441\u043d\u0438\u043c\u043e\u043a'],\n  provider_failure: ['\u043f\u043e\u0432\u0442\u043e\u0440\u0438\u0442\u044c \u043f\u043e\u043f\u044b\u0442\u043a\u0443', '\u043f\u0440\u043e\u0432\u0435\u0440\u0438\u0442\u044c \u0434\u0440\u0443\u0433\u043e\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a'],\n  timeout: ['\u043f\u043e\u0432\u0442\u043e\u0440\u0438\u0442\u044c \u043f\u043e\u043f\u044b\u0442\u043a\u0443', '\u043f\u0440\u043e\u0432\u0435\u0440\u0438\u0442\u044c \u0434\u0440\u0443\u0433\u043e\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a'],\n  empty_response: ['\u0443\u043a\u0430\u0437\u0430\u0442\u044c \u043a\u043e\u043d\u043a\u0440\u0435\u0442\u043d\u0443\u044e \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0443 \u0443\u0441\u043b\u0443\u0433\u0438', '\u043f\u0440\u043e\u0432\u0435\u0440\u0438\u0442\u044c \u0434\u0440\u0443\u0433\u043e\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a'],\n  unsupported_content: ['\u0443\u043a\u0430\u0437\u0430\u0442\u044c \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0443 \u0441 \u0442\u0435\u043a\u0441\u0442\u043e\u043c (\u043d\u0430\u043f\u0440\u0438\u043c\u0435\u0440, \u0440\u0430\u0437\u0434\u0435\u043b \u0443\u0441\u043b\u0443\u0433)']\n};\nfunction saNextActionsRu(outcome, opts) {\n  var list = (SA_NEXT_RU[saStr(outcome)] || ['\u043f\u043e\u0432\u0442\u043e\u0440\u0438\u0442\u044c \u043f\u043e\u043f\u044b\u0442\u043a\u0443 \u043f\u043e\u0437\u0436\u0435']).slice();\n  // Only offer the saved snapshot when one actually exists.\n  if (!(opts && opts.has_snapshot)) list = list.filter(function (x) { return x.indexOf('\u0441\u043e\u0445\u0440\u0430\u043d\u0451\u043d\u043d') < 0; });\n  return list.slice(0, 3);\n}\n// --- end embedded source_access ---\n\nconst resp = $json;\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nconst ctx = $('Evaluate Dedup').first().json;\nconst targetUrl = ctx.target_url || ctx.source_url || '';\nconst sourceUrlBase = ctx.source_url || targetUrl;\nconst run_id = ctx.run_id || '';\nconst batch_index = ctx.batch_index || 0;\nconst parsedAt = ctx.parsed_at || moscowIsoNow();\nconst now = moscowIsoNow();\nfunction cap(s, n) { return (s == null ? '' : String(s)).substring(0, n); }\n\nfunction technicalErrorRow(errSummary, preview) {\n  return [{ json: {\n    created_at: now, source_type: 'scraped_web', platform: 'website', source_url: sourceUrlBase, parsed_at: parsedAt,\n    published_at: '', freshness_status: 'unknown', entity_type: 'irrelevant', company_name: '', profile_name: '',\n    profile_url: '', region: '', service_type: 'unknown', offer_text: '', terms: '', contact_public: '',\n    text_context: '', detected_need: '', competitor_strength: 1, lead_signal_score: 1, content_idea_score: 1,\n    quality_score: 1, reason: '', recommended_action: 'ignore', status: 'skipped',\n    processing_status: 'technical_error', parse_method: 'firecrawl_error',\n    parse_error: cap('Firecrawl scrape failed: ' + errSummary, 800),\n    raw_response_preview: cap(preview, 500), route: 'technical_errors', needs_manual_review: true,\n    repair_used: false, repair_status: '', run_id: run_id, batch_index: batch_index\n  }}];\n}\n\nconst apiError = (resp == null) || resp.error || resp.success === false || (resp.code && resp.code >= 400) || (resp.statusCode && resp.statusCode >= 400);\nif (apiError) {\n  const summary = cap(JSON.stringify((resp && (resp.error || resp.message || resp.code || resp.statusCode)) || 'unknown error'), 300);\n  return technicalErrorRow(summary, cap(JSON.stringify(resp), 500));\n}\n\nconst data = resp.data || resp;\nlet markdown = (resp.data && resp.data.markdown) || resp.markdown || (data && data.markdown) || (resp.data && resp.data.data && resp.data.data.markdown) || '';\nconst metadata = (resp.data && resp.data.metadata) || resp.metadata || (data && data.metadata) || {};\nmarkdown = String(markdown).replace(/\\r/g, '');\n\n// Clean: drop image lines and inline image/svg data fragments\nconst cleanedLines = [];\nfor (const line of markdown.split('\\n')) {\n  const t = line.trim();\n  if (t.startsWith('![')) continue;\n  if (/data:image\\/svg\\+xml/i.test(line)) continue;\n  if (/data:image\\/[a-z]+;base64/i.test(line)) continue;\n  cleanedLines.push(line);\n}\nlet cleaned = cleanedLines.join('\\n').replace(/\\n{3,}/g, '\\n\\n').trim();\n\nconst meaningful = cleaned.replace(/[#>*_|]/g, '').replace(/\\s+/g, ' ').trim();\n// BLOCK-HONESTY-001: decide whether we REACHED the page BEFORE anyone asks whether the business is relevant.\n// Live (carmoney.ru, exec 894): a 772-char access-restriction page cleared the 80-char floor, reached Claude, and\n// came back \"irrelevant\" \u2014 so the user was told the site was checked and found irrelevant. We never saw the site.\n// An access failure is NOT a business verdict: it must never become a competitor snapshot, never overwrite a good\n// snapshot, never reach Claude (that also saves the call), and never be reported as \"\u043f\u0440\u043e\u0432\u0435\u0440\u0435\u043d\".\nconst __access = classifySourceAccess({ status: 200, body_text: cleaned, content_type: 'text/html' });\nif (__access.access_failure) {\n  const __row = technicalErrorRow('source not accessible: ' + __access.outcome + ' (' + __access.reason + ')', cleaned)[0];\n  // entity_type stays 'unknown' \u2014 NOT 'irrelevant'. We never judged the business.\n  __row.json.entity_type = 'unknown';\n  __row.json.parse_method = 'source_' + __access.outcome;\n  __row.json.access_outcome = __access.outcome;\n  __row.json.access_retryable = __access.retryable === true;\n  __row.json.reason = saUserMessageRu(__access.outcome);\n  __row.json.recommended_action = 'ignore';\n  return [__row];\n}\nif (!cleaned || meaningful.length < 80) {\n  return technicalErrorRow('scrape succeeded but markdown empty/unusable (' + meaningful.length + ' meaningful chars)', cap(cleaned || JSON.stringify(metadata), 500));\n}\n\n// Commercially relevant lines first (stable order within groups), then cap at 3500\nconst relevantTerms = ['\u043a\u0440\u0435\u0434\u0438\u0442','\u0437\u0430\u0439\u043c','\u0437\u0430\u043b\u043e\u0433','\u043f\u0442\u0441','\u0430\u0432\u0442\u043e','\u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b','\u043d\u0435\u0434\u0432\u0438\u0436','\u043a\u0432\u0430\u0440\u0442\u0438\u0440','\u0434\u043e\u043c','\u0437\u0435\u043c\u043b','\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441','\u0438\u043f\u043e\u0442\u0435\u043a','\u0441\u0442\u0430\u0432\u043a\u0430','\u0441\u0443\u043c\u043c\u0430','\u043e\u0434\u043e\u0431\u0440\u0435\u043d','\u043f\u043b\u043e\u0445\u0430\u044f \u043a\u0440\u0435\u0434\u0438\u0442\u043d','\u043f\u0440\u043e\u0441\u0440\u043e\u0447','\u0442\u0435\u043b\u0435\u0444\u043e\u043d','\u043c\u043e\u0441\u043a\u0432\u0430','\u0440\u0443\u0431','%'];\nconst relevant = [];\nconst rest = [];\nfor (const line of cleaned.split('\\n')) {\n  const low = line.toLowerCase();\n  if (relevantTerms.some(t => low.includes(t))) relevant.push(line); else rest.push(line);\n}\nlet prioritized = (relevant.length > 0) ? (relevant.join('\\n') + '\\n' + rest.join('\\n')).replace(/\\n{3,}/g, '\\n\\n').trim() : cleaned;\n\n// Placeholder / parking / domain-not-connected pre-filter (DEC-054): skip BEFORE Claude cost.\n// Strong phrases skip unconditionally; bare 'coming soon' only when the page has no business content.\nconst placeholderHay = (cleaned + ' ' + (metadata.title || '') + ' ' + (metadata.description || '')).toLowerCase();\nconst strongPlaceholder = ['domain is not connected','domain not connected','not connected to any website','wix domain not connected','this domain is not connected','parking page','\u0441\u0430\u0439\u0442 \u043d\u0435 \u043f\u043e\u0434\u043a\u043b\u044e\u0447\u0435\u043d','\u0434\u043e\u043c\u0435\u043d \u043d\u0435 \u043f\u043e\u0434\u043a\u043b\u044e\u0447\u0435\u043d','\u0437\u0430\u0433\u043b\u0443\u0448\u043a\u0430 \u0441\u0430\u0439\u0442\u0430'];\nconst isStrongPlaceholder = strongPlaceholder.some(s => placeholderHay.includes(s));\nconst isComingSoonEmpty = placeholderHay.includes('coming soon') && relevant.length === 0;\nif (isStrongPlaceholder || isComingSoonEmpty) {\n  return [{ json: {\n    created_at: now, source_type: 'scraped_web', platform: 'website', source_url: sourceUrlBase, parsed_at: parsedAt,\n    published_at: '', freshness_status: 'unknown', entity_type: 'irrelevant', company_name: '', profile_name: '',\n    profile_url: '', region: '', service_type: 'unknown', offer_text: '', terms: '', contact_public: '',\n    text_context: '', detected_need: '', competitor_strength: 1, lead_signal_score: 1, content_idea_score: 1,\n    quality_score: 1, reason: 'Firecrawl returned a placeholder/domain-not-connected page; skipped before Claude cost.',\n    recommended_action: 'ignore', status: 'skipped', processing_status: 'business_skip',\n    parse_method: 'firecrawl_placeholder_prefilter', parse_error: '',\n    raw_response_preview: cap(cleaned, 500), route: 'skipped_log', needs_manual_review: false,\n    repair_used: false, repair_status: '', run_id: run_id, batch_index: batch_index\n  }}];\n}\n\nconst sourceUrl = metadata.sourceURL || metadata.url || sourceUrlBase;\nconst title = metadata.title || '';\nconst description = metadata.description || '';\n\nreturn [{ json: {\n  route: '', source_type: 'scraped_web', platform: 'website', source_url: sourceUrl, profile_url: '',\n  published_at: '', parsed_at: parsedAt, text_context: cap(prioritized, 3500),\n  page_title: cap(title, 300), page_description: cap(description, 500),\n  run_id: run_id, batch_index: batch_index\n}}];"
      }
    },
    {
      "id": "fc000000-0000-0000-0000-000000000005",
      "name": "IF Firecrawl Normalized OK?",
      "type": "n8n-nodes-base.if",
      "typeVersion": 2,
      "position": [
        2200,
        300
      ],
      "parameters": {
        "conditions": {
          "options": {
            "caseSensitive": true,
            "leftValue": "",
            "typeValidation": "loose"
          },
          "conditions": [
            {
              "id": "fc-if-01",
              "leftValue": "={{ $json.route }}",
              "rightValue": "",
              "operator": {
                "type": "string",
                "operation": "empty",
                "singleValue": true
              }
            }
          ],
          "combinator": "and"
        }
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000005",
      "name": "Build Primary Claude Request",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        2420,
        140
      ],
      "parameters": {
        "jsCode": "const record = $json;\n\nconst systemPrompt = `You are Marketing Scout Agent v2 -- a market intelligence analyst for a secured lending business in Moscow and Moscow Oblast, Russia.\n\nFor every record ask: What does this mean for the operator's business and what should they do? Reason like a business owner.\n\nANALYSIS PRIORITY ORDER:\n1. LEAD SIGNAL first -- potential client needing secured loan? Moscow/MO? PTS/auto/real estate collateral? Urgency? Contactable?\n2. COMPETITOR second -- active secured lending business Moscow/MO? Threat level?\n3. CONTENT IDEA third -- client fear/objection/knowledge gap for secured lending?\n4. IRRELEVANT if none apply.\n\nIDEAL CLIENT: Car owner (PTS clean title), Moscow/MO, needs cash urgently, bank-rejected, 50k-500k RUB.\n\nHigh-urgency signals (raise lead_signal_score): \u0441\u0440\u043e\u0447\u043d\u043e, \u0441\u0435\u0433\u043e\u0434\u043d\u044f, \u0431\u0430\u043d\u043a\u0438 \u043e\u0442\u043a\u0430\u0437\u0430\u043b\u0438, \u043d\u0435 \u0434\u0430\u044e\u0442 \u043a\u0440\u0435\u0434\u0438\u0442, \u0438\u0441\u043f\u043e\u0440\u0447\u0435\u043d\u0430 \u043a\u0440\u0435\u0434\u0438\u0442\u043d\u0430\u044f \u0438\u0441\u0442\u043e\u0440\u0438\u044f + specific amount + collateral type.\n\nREGION RULES:\n- Moscow/MO explicit: lead_signal_score 60-100\n- Region ambiguous: eligible up to 55\n- Another city/region: lead_signal_score capped at 40\nCompetitors: Moscow/MO or national coverage -> score normally; other region only -> cap competitor_strength at 50.\n\nlead_signal_score calibration:\n- 85-100: fit + urgency + readiness + Moscow/MO confirmed\n- 70-84: strong fit+urgency, region confirmed, readiness partial\n- 55-69: clear product fit, one signal confirmed, region present\n- 35-54: intent apparent, fit unclear or region outside MO\n- 1-34: no real lead signal\n\nrecommended_action=contact requires lead_signal_score>=70.\n\ncompetitor_strength calibration:\n- 85-100: fresh (<=30d), Moscow/MO, stated rate, same-day, bad-credit accepted, contactable\n- 65-84: active professional, Moscow/MO confirmed, rate absent or one signal missing\n- 45-64: present but older content or uncertain coverage\n- 25-44: weak - stale or different region\n- 1: not a competitor\n\nSKIP RULES -- return status=skipped, quality_score=1, all scores=1 when:\n- Fewer than 40 meaningful chars\n- Pure navigation boilerplate\n- No connection to financial services\n- published_at >180 days before parsed_at with no fresh signals\n\nREASON FIELD (3 sentences required):\n1. WHAT: what is this record, key evidence from text\n2. WHY: why scores are what they are, cite specific signals\n3. NEXT: what operator should do and why\n\nOUTPUT FORMAT -- CRITICAL:\nRespond with ONLY a valid JSON object.\nNo markdown, no code fences, no preamble.\nFirst character must be {. Last must be }.\nAll 25 fields required. Empty string for unknown strings. 1 for unknown integers.\n\nREQUIRED JSON SCHEMA:\n{\n  \"created_at\": \"<parsed_at value ISO 8601>\",\n  \"source_type\": \"<from input>\",\n  \"platform\": \"<from input>\",\n  \"source_url\": \"<from input>\",\n  \"parsed_at\": \"<from input>\",\n  \"published_at\": \"<from input or empty>\",\n  \"freshness_status\": \"<fresh|recent|old|unknown>\",\n  \"entity_type\": \"<competitor|lead_signal|market_signal|content_idea|irrelevant>\",\n  \"company_name\": \"<explicitly in text only or empty>\",\n  \"profile_name\": \"<explicitly in text only or empty>\",\n  \"profile_url\": \"<from input or empty>\",\n  \"region\": \"<explicitly mentioned or empty>\",\n  \"service_type\": \"<secured_auto_loan|secured_real_estate_loan|pts_loan|refinancing|mortgage_adjacent|generic_lending|unknown>\",\n  \"offer_text\": \"<1 sentence: what offered/sought or content angle title>\",\n  \"terms\": \"<explicit rate/conditions only or empty>\",\n  \"contact_public\": \"<phone/email/Telegram from text only or empty>\",\n  \"text_context\": \"<cleaned summary max 300 chars>\",\n  \"detected_need\": \"<lead_signal only: need+amount+urgency+bank rejection+region or empty>\",\n  \"competitor_strength\": <integer 1-100; 1 if not competitor>,\n  \"lead_signal_score\": <integer 1-100>,\n  \"content_idea_score\": <integer 1-100>,\n  \"quality_score\": <integer 1-100>,\n  \"reason\": \"<3 sentences: what+evidence; why scores; next action>\",\n  \"recommended_action\": \"<monitor|contact|create_content|ignore|investigate>\",\n  \"status\": \"<analyzed|skipped>\"\n}\n\nREMINDER: Return JSON only. No markdown. No analysis outside JSON. For competitor website records, classify entity_type=competitor if the text offers secured lending services, rates, speed, contact, or Moscow/MO coverage.`;\n\nreturn [{ json: {\n  model: 'claude-sonnet-4-6',\n  max_tokens: 1400,\n  temperature: 0.2,\n  system: systemPrompt,\n  messages: [{\n    role: 'user',\n    content: JSON.stringify({\n      source_type: record.source_type || '',\n      platform: record.platform || '',\n      source_url: record.source_url || '',\n      profile_url: record.profile_url || '',\n      parsed_at: record.parsed_at || '',\n      published_at: record.published_at || '',\n      text_context: record.text_context || ''\n    })\n  }]\n}}];"
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000007",
      "name": "Claude Primary API Request",
      "type": "n8n-nodes-base.httpRequest",
      "typeVersion": 4.2,
      "position": [
        2640,
        140
      ],
      "onError": "continueRegularOutput",
      "parameters": {
        "method": "POST",
        "url": "https://aiprimetech.io/v1/messages",
        "authentication": "predefinedCredentialType",
        "nodeCredentialType": "httpHeaderAuth",
        "sendHeaders": true,
        "headerParameters": {
          "parameters": [
            {
              "name": "anthropic-version",
              "value": "2023-06-01"
            }
          ]
        },
        "sendBody": true,
        "contentType": "raw",
        "rawContentType": "application/json",
        "body": "={{ JSON.stringify({ model: $json.model, max_tokens: $json.max_tokens, temperature: $json.temperature, system: $json.system, messages: $json.messages }) }}",
        "options": {}
      },
      "credentials": {
        "httpHeaderAuth": {
          "name": "<your credential>"
        }
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000008",
      "name": "Parse Primary JSON",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        2860,
        140
      ],
      "parameters": {
        "jsCode": "const response = $json;\nconst srcRecord = $('Normalize Firecrawl Output').first().json;\nconst __sd=$getWorkflowStaticData('global');const __r=(__sd.wf04_run=__sd.wf04_run||{});__r.primary_calls=(__r.primary_calls||0)+1;__r.firecrawl_calls=(__r.firecrawl_calls||0)+1;__r.urls_scraped=(__r.urls_scraped||0)+1;__r.claude_calls=(__r.claude_calls||0)+1;function __pp(ok){if(ok)__r.primary_parse_success=(__r.primary_parse_success||0)+1;else __r.primary_parse_failure=(__r.primary_parse_failure||0)+1;}\nfunction cap(s, n) { return (s == null ? '' : String(s)).substring(0, n); }\nfunction extractBalanced(s) {\n  const start = s.indexOf('{');\n  if (start === -1) return null;\n  let depth = 0, inStr = false, esc = false;\n  for (let i = start; i < s.length; i++) {\n    const c = s[i];\n    if (inStr) { if (esc) { esc = false; } else if (c === '\\\\') { esc = true; } else if (c === '\"') { inStr = false; } continue; }\n    if (c === '\"') { inStr = true; continue; }\n    if (c === '{') depth++;\n    else if (c === '}') { depth--; if (depth === 0) return s.substring(start, i + 1); }\n  }\n  return null;\n}\nfunction tryParse(text) {\n  let txt = String(text).trim();\n  txt = txt.replace(/^```json\\s*/i, '').replace(/```\\s*$/, '').trim();\n  txt = txt.replace(/^```\\s*/, '').replace(/```\\s*$/, '').trim();\n  const norm = (x) => x.replace(/[\u2018\u2019]/g, \"'\").replace(/[\u201c\u201d\u00ab\u00bb]/g, '\"');\n  try { return { ok: true, obj: JSON.parse(norm(txt)), candidate: cap(txt, 500) }; } catch(e) {}\n  const bal = extractBalanced(txt);\n  if (bal) { try { return { ok: true, obj: JSON.parse(norm(bal)), candidate: cap(bal, 500) }; } catch(e) {} }\n  const s = txt.indexOf('{'), e2 = txt.lastIndexOf('}');\n  if (s !== -1 && e2 !== -1 && e2 > s) { const slice = txt.substring(s, e2 + 1); try { return { ok: true, obj: JSON.parse(norm(slice)), candidate: cap(slice, 500) }; } catch(e) {} return { ok: false, candidate: cap(slice, 500) }; }\n  return { ok: false, candidate: cap(txt, 500) };\n}\nfunction fail(err, prev, summary) {\n  __pp(false);\n  return [{ json: {\n    parse_ok: false, parse_method: 'primary_json', repair_used: false,\n    processing_status: 'technical_error_candidate',\n    primary_parse_error: cap(err, 300),\n    primary_raw_response_preview: cap(prev, 500),\n    extracted_candidate_preview: cap(prev, 500),\n    content_summary: summary,\n    parse_error: cap(err, 300),\n    raw_response_preview: cap(prev, 500),\n    original_record: srcRecord\n  }}];\n}\n\nif (response.error || response.status >= 400) {\n  const prev = cap(JSON.stringify(response), 500);\n  return fail('Primary HTTP error: ' + cap(JSON.stringify(response.error || response.status), 300), prev, 'http_error');\n}\n\nlet textItem, contentSummary = '';\ntry {\n  if (!response.content || !Array.isArray(response.content)) throw new Error('no content array');\n  contentSummary = response.content.map(c => c.type).join(',');\n  textItem = response.content.find(c => c.type === 'text');\n} catch(e) {\n  return fail('Primary parse failed: no content array: ' + e.message, cap(JSON.stringify(response), 500), 'no_content_array');\n}\nif (!textItem || !textItem.text) {\n  return fail('Primary parse failed: no text item', cap(JSON.stringify(response.content || response), 500), contentSummary || 'no_text_item');\n}\n\nconst rawPreview = cap(textItem.text, 500);\nconst res = tryParse(textItem.text);\nif (!res.ok) {\n  const out = fail('Primary JSON parse failed', rawPreview, 'text_present_json_invalid');\n  out[0].json.extracted_candidate_preview = cap(res.candidate || '', 500);\n  return out;\n}\n\n__pp(true);\nreturn [{ json: {\n  parse_ok: true,\n  parse_method: 'primary_json',\n  repair_used: false,\n  repair_status: '',\n  parse_error: '',\n  raw_response_preview: rawPreview,\n  ...res.obj\n}}];"
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000009",
      "name": "IF Primary Parse OK?",
      "type": "n8n-nodes-base.if",
      "typeVersion": 2,
      "position": [
        3080,
        140
      ],
      "parameters": {
        "conditions": {
          "options": {
            "caseSensitive": true,
            "leftValue": "",
            "typeValidation": "loose"
          },
          "conditions": [
            {
              "id": "rr000000-0000-0000-0000-000000000091",
              "leftValue": "={{ $json.parse_ok }}",
              "rightValue": true,
              "operator": {
                "type": "boolean",
                "operation": "true"
              }
            }
          ],
          "combinator": "and"
        }
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000010",
      "name": "Build Repair Request",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        3080,
        560
      ],
      "parameters": {
        "jsCode": "const input = $json;\nconst srcRecord = $('Normalize Firecrawl Output').first().json;\nfunction cap(s, n) { return (s == null ? '' : String(s)).substring(0, n); }\n\nconst rawPreview = cap(input.primary_raw_response_preview || input.raw_response_preview || 'empty', 500);\nconst primaryErr = cap(input.primary_parse_error || input.parse_error || 'Primary parse failed', 300);\n\nconst repairSystem = `You are a JSON repair formatter, not a market analyst. Convert the raw response into strict JSON. Do not invent facts. If raw response is unusable, return an object that can be routed to technical_errors.\nReturn ONLY a valid JSON object. Do not say Repairing. Do not explain. Do not use markdown.\nReturn JSON only. No markdown. First char {, last char }.\nFields: created_at, source_type, platform, source_url, parsed_at, published_at, freshness_status, entity_type, company_name, profile_name, profile_url, region, service_type, offer_text, terms, contact_public, text_context, detected_need, competitor_strength, lead_signal_score, content_idea_score, quality_score, reason, recommended_action, status.\nUnknown text -> \"\". Unknown score -> 1. Unusable -> status=skipped, entity_type=irrelevant, recommended_action=ignore, all scores=1.\nEnums: entity_type[competitor,lead_signal,market_signal,content_idea,irrelevant]; recommended_action[monitor,contact,create_content,ignore,investigate]; status[analyzed,skipped]; freshness_status[fresh,recent,old,unknown].`;\n\nconst userMsg = JSON.stringify({\n  original_record: {\n    source_type: srcRecord.source_type || '',\n    platform: srcRecord.platform || '',\n    source_url: srcRecord.source_url || '',\n    profile_url: srcRecord.profile_url || '',\n    parsed_at: srcRecord.parsed_at || '',\n    published_at: srcRecord.published_at || '',\n    text_context: cap(srcRecord.text_context || '', 500)\n  },\n  raw_primary_response: rawPreview,\n  primary_parse_error: primaryErr\n});\n\nreturn [{ json: {\n  model: 'claude-sonnet-4-6',\n  max_tokens: 700,\n  temperature: 0,\n  system: repairSystem,\n  messages: [{ role: 'user', content: userMsg }]\n}}];"
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000011",
      "name": "Claude Repair API Request",
      "type": "n8n-nodes-base.httpRequest",
      "typeVersion": 4.2,
      "position": [
        3300,
        560
      ],
      "onError": "continueRegularOutput",
      "parameters": {
        "method": "POST",
        "url": "https://aiprimetech.io/v1/messages",
        "authentication": "predefinedCredentialType",
        "nodeCredentialType": "httpHeaderAuth",
        "sendHeaders": true,
        "headerParameters": {
          "parameters": [
            {
              "name": "anthropic-version",
              "value": "2023-06-01"
            }
          ]
        },
        "sendBody": true,
        "contentType": "raw",
        "rawContentType": "application/json",
        "body": "={{ JSON.stringify({ model: $json.model, max_tokens: $json.max_tokens, temperature: $json.temperature, system: $json.system, messages: $json.messages }) }}",
        "options": {}
      },
      "credentials": {
        "httpHeaderAuth": {
          "name": "<your credential>"
        }
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000012",
      "name": "Parse Repaired JSON",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        3520,
        560
      ],
      "parameters": {
        "jsCode": "const response = $json;\nconst __sd=$getWorkflowStaticData('global');const __r=(__sd.wf04_run=__sd.wf04_run||{});__r.repair_calls=(__r.repair_calls||0)+1;__r.claude_calls=(__r.claude_calls||0)+1;\nconst primary = $('Parse Primary JSON').first().json;\nconst src = $('Normalize Firecrawl Output').first().json;\nfunction cap(s, n) { return (s == null ? '' : String(s)).substring(0, n); }\nfunction extractBalanced(s) {\n  const start = s.indexOf('{');\n  if (start === -1) return null;\n  let depth = 0, inStr = false, esc = false;\n  for (let i = start; i < s.length; i++) {\n    const c = s[i];\n    if (inStr) { if (esc) { esc = false; } else if (c === '\\\\') { esc = true; } else if (c === '\"') { inStr = false; } continue; }\n    if (c === '\"') { inStr = true; continue; }\n    if (c === '{') depth++;\n    else if (c === '}') { depth--; if (depth === 0) return s.substring(start, i + 1); }\n  }\n  return null;\n}\nfunction tryParse(text) {\n  let txt = String(text).trim();\n  txt = txt.replace(/^```json\\s*/i, '').replace(/```\\s*$/, '').trim();\n  txt = txt.replace(/^```\\s*/, '').replace(/```\\s*$/, '').trim();\n  const norm = (x) => x.replace(/[\u2018\u2019]/g, \"'\").replace(/[\u201c\u201d\u00ab\u00bb]/g, '\"');\n  try { return { ok: true, obj: JSON.parse(norm(txt)) }; } catch(e) {}\n  const bal = extractBalanced(txt);\n  if (bal) { try { return { ok: true, obj: JSON.parse(norm(bal)) }; } catch(e) {} }\n  const s = txt.indexOf('{'), e2 = txt.lastIndexOf('}');\n  if (s !== -1 && e2 !== -1 && e2 > s) { try { return { ok: true, obj: JSON.parse(norm(txt.substring(s, e2 + 1))) }; } catch(e) {} }\n  return { ok: false };\n}\n\nconst primaryErr = cap(primary.primary_parse_error || primary.parse_error || 'unknown', 300);\nconst primaryPreview = cap(primary.primary_raw_response_preview || primary.raw_response_preview || '', 500);\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nconst now = moscowIsoNow();\nconst run_id = src.run_id || '';\nconst batch_index = src.batch_index || 0;\nconst parsedAt = src.parsed_at || now;\n\n// Deterministic competitor fallback after primary + repair failure (DEC-052).\nfunction fallback(repairErr) {\n  const combinedErr = cap('Primary: ' + primaryErr + ' | Repair: ' + repairErr, 800);\n  const textCtx = String(src.text_context || '');\n  const rawAll = (textCtx + ' ' + primaryPreview).toLowerCase();\n  const signals = ['\u043a\u0440\u0435\u0434\u0438\u0442','\u0437\u0430\u0439\u043c','\u0437\u0430\u043b\u043e\u0433','\u043f\u0442\u0441','\u0430\u0432\u0442\u043e','\u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b','\u043d\u0435\u0434\u0432\u0438\u0436','\u043a\u0432\u0430\u0440\u0442\u0438\u0440','\u0434\u043e\u043c','\u0437\u0435\u043c\u043b','\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441','\u0438\u043f\u043e\u0442\u0435\u043a','\u0441\u0442\u0430\u0432\u043a\u0430','\u0441\u0443\u043c\u043c\u0430','\u043e\u0434\u043e\u0431\u0440\u0435\u043d','\u043f\u043b\u043e\u0445\u0430\u044f \u043a\u0440\u0435\u0434\u0438\u0442\u043d','\u043f\u0440\u043e\u0441\u0440\u043e\u0447','\u0442\u0435\u043b\u0435\u0444\u043e\u043d','\u043c\u043e\u0441\u043a\u0432\u0430','\u0440\u0443\u0431','%'];\n  let count = 0;\n  for (const t of signals) { if (rawAll.includes(t)) count++; }\n\n  if (count < 5) {\n    let preview = primaryPreview;\n    if (preview.length < 440) preview = cap(preview + ' || REPAIR_ERR: ' + repairErr, 500);\n    return [{ json: {\n      parse_ok: false, parse_method: 'technical_error', repair_used: true, repair_status: 'failed',\n      processing_status: 'technical_error', route: 'technical_errors', needs_manual_review: true,\n      parse_error: combinedErr, raw_response_preview: cap(preview, 500),\n      original_record: primary.original_record || {}, run_id: run_id, batch_index: batch_index\n    }}];\n  }\n\n  const low = textCtx.toLowerCase();\n  let company = '\u041a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442 \u0431\u0435\u0437 \u0431\u0440\u0435\u043d\u0434\u0430';\n  if (low.includes('\u043c\u043e\u0441\u0438\u043d\u0432\u0435\u0441\u0442\u0444\u0438\u043d\u0430\u043d\u0441') || low.includes('mosinvest')) company = '\u041c\u043e\u0441\u0418\u043d\u0432\u0435\u0441\u0442\u0424\u0438\u043d\u0430\u043d\u0441';\n  else if (low.includes('lioncredit') || low.includes('lion credit')) company = 'LionCredit';\n  else { try { const h = new URL(src.source_url || '').hostname.replace(/^www\\./, ''); if (h) company = h; } catch(e) {} }\n\n  const hasAuto = low.includes('\u043f\u0442\u0441') || low.includes('pts') || low.includes('\u0430\u0432\u0442\u043e') || low.includes('\u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b') || low.includes('\u043c\u0430\u0448\u0438\u043d');\n  const hasRealEstate = low.includes('\u043d\u0435\u0434\u0432\u0438\u0436') || low.includes('\u043a\u0432\u0430\u0440\u0442\u0438\u0440') || low.includes('\u0434\u043e\u043c') || low.includes('\u0437\u0435\u043c\u043b');\n  let service = 'generic_lending';\n  if (hasAuto && !hasRealEstate) service = 'pts_loan';\n  else if (hasRealEstate && !hasAuto) service = 'secured_real_estate_loan';\n\n  let offer = '';\n  for (const line of textCtx.split('\\n')) { const l = line.trim().toLowerCase(); if ((l.includes('\u043a\u0440\u0435\u0434\u0438\u0442') || l.includes('\u0437\u0430\u0439\u043c') || l.includes('\u0437\u0430\u043b\u043e\u0433')) && line.trim().length > 10) { offer = line.trim(); break; } }\n  offer = cap(offer, 220);\n\n  let terms = '';\n  for (const line of textCtx.split('\\n')) { const l = line.toLowerCase(); if (l.includes('%') || l.includes('\u0441\u0442\u0430\u0432\u043a') || l.includes('\u0441\u0443\u043c\u043c\u0430') || l.includes('\u0440\u0443\u0431')) { terms += (terms ? ' ' : '') + line.trim(); if (terms.length >= 200) break; } }\n  terms = cap(terms, 220);\n\n  const region = (low.includes('\u043c\u043e\u0441\u043a\u0432\u0430') || low.includes('\u043c\u043e\u0441\u043a\u043e\u0432\u0441\u043a')) ? '\u041c\u043e\u0441\u043a\u0432\u0430' : '';\n  const hasContactOrAmount = low.includes('\u0442\u0435\u043b\u0435\u0444\u043e\u043d') || /\\+?\\d[\\d\\s().-]{7,}/.test(textCtx) || low.includes('\u0440\u0443\u0431') || low.includes('%');\n  const strength = (count >= 8 && hasContactOrAmount) ? 80 : 70;\n\n  return [{ json: {\n    created_at: now, source_type: 'scraped_web', platform: 'website', source_url: src.source_url || '', parsed_at: parsedAt,\n    published_at: '', freshness_status: 'unknown', entity_type: 'competitor', company_name: company,\n    profile_name: '', profile_url: '', region: region, service_type: service,\n    offer_text: offer, terms: terms, contact_public: '', text_context: cap(textCtx, 3500), detected_need: '',\n    competitor_strength: strength, lead_signal_score: 1, content_idea_score: 1, quality_score: strength,\n    reason: 'Claude \u0432\u0435\u0440\u043d\u0443\u043b \u043d\u0435\u0432\u0430\u043b\u0438\u0434\u043d\u044b\u0439 JSON, \u043d\u043e \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u0430 \u0441\u043e\u0434\u0435\u0440\u0436\u0438\u0442 \u044f\u0432\u043d\u044b\u0435 \u043f\u0440\u0438\u0437\u043d\u0430\u043a\u0438 \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u0430 \u0432 \u043d\u0438\u0448\u0435 \u0437\u0430\u043b\u043e\u0433\u043e\u0432\u043e\u0433\u043e \u043a\u0440\u0435\u0434\u0438\u0442\u043e\u0432\u0430\u043d\u0438\u044f: \u043f\u0440\u043e\u0434\u0443\u043a\u0442\u044b, \u0443\u0441\u043b\u043e\u0432\u0438\u044f, \u0441\u0443\u043c\u043c\u044b/\u0441\u0442\u0430\u0432\u043a\u0438 \u0438\u043b\u0438 \u043a\u043e\u043d\u0442\u0430\u043a\u0442. \u0417\u0430\u043f\u0438\u0441\u044c \u043e\u0442\u043f\u0440\u0430\u0432\u043b\u0435\u043d\u0430 \u0432 \u043c\u043e\u043d\u0438\u0442\u043e\u0440\u0438\u043d\u0433, \u0447\u0442\u043e\u0431\u044b \u043d\u0435 \u0442\u0435\u0440\u044f\u0442\u044c \u043f\u043e\u043b\u0435\u0437\u043d\u044b\u0439 \u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a.',\n    recommended_action: 'monitor', status: 'analyzed', processing_status: 'parsed_success',\n    parse_method: 'deterministic_competitor_fallback', parse_error: combinedErr,\n    raw_response_preview: cap(primaryPreview, 500), route: 'monitor_queue', needs_manual_review: true,\n    repair_used: true, repair_status: 'failed_fallback', run_id: run_id, batch_index: batch_index\n  }}];\n}\n\nif (response.error || response.status >= 400) {\n  return fallback('Repair HTTP error: ' + cap(JSON.stringify(response.error || response.status), 300));\n}\n\nlet textItem;\ntry { textItem = (response.content || []).find(c => c.type === 'text'); } catch(e) { textItem = null; }\nif (!textItem || !textItem.text) {\n  return fallback('No text item in repair response');\n}\n\nconst res = tryParse(textItem.text);\nif (!res.ok) {\n  return fallback('Repair JSON parse failed');\n}\n\nreturn [{ json: {\n  parse_ok: true,\n  parse_method: 'repaired_json',\n  repair_used: true,\n  repair_status: 'success',\n  parse_error: '',\n  raw_response_preview: cap(textItem.text, 500),\n  ...res.obj\n}}];"
      }
    },
    {
      "id": "rr000000-0000-0000-0000-000000000013",
      "name": "Normalize + Route",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        3740,
        300
      ],
      "parameters": {
        "jsCode": "// embedded n8n/lib/error_sanitizer.js (drift-proof; test asserts equality)\n// error_sanitizer.js \u2014 what may be persisted about a FAILED record (technical_errors / skipped_log routes).\n//\n// WF04-ROUTE-002. The resilient routers persist `raw_response_preview` + `parse_error` for triage. Those strings are\n// built from a provider response, so without a gate they can carry an Authorization header, an api key, a cookie, a\n// Claude thinking block, or a customer's phone/email straight into a durable Google Sheets tab that many people can\n// open. Diagnosis needs a bounded, scrubbed EXCERPT \u2014 never the raw body.\n//\n// Keep: provider, source URL, safe error category, bounded sanitized excerpt, request/run lineage, timestamp.\n// Drop: secrets, credentials, cookies, hidden reasoning, private PII, anything past the cap.\n//\n// Embeddable: unique es*-prefixed names, no cross-lib require.\n\nfunction esStr(v) { return v == null ? '' : String(v); }\n\nvar ES_MAX_PREVIEW = 300;      // enough to recognise a failure shape; far too short to be a \"raw body\"\nvar ES_REDACTED = '[\u0441\u043a\u0440\u044b\u0442\u043e]';\n\n// Ordered: the most specific secret shapes first, then generic key/value pairs, then PII.\nvar ES_RULES = [\n  // Authorization / bearer / api-key headers (with or without a header name)\n  { re: /(authorization|proxy-authorization)\\s*[:=]\\s*\\S+/gi, to: '$1: ' + ES_REDACTED },\n  { re: /\\bbearer\\s+[A-Za-z0-9._\\-~+/]{8,}=*/gi, to: 'bearer ' + ES_REDACTED },\n  { re: /\\bbasic\\s+[A-Za-z0-9+/]{8,}=*/gi, to: 'basic ' + ES_REDACTED },\n  // cookies / set-cookie\n  { re: /(set-cookie|cookie)\\s*[:=]\\s*[^\\n;]+/gi, to: '$1: ' + ES_REDACTED },\n  // provider key formats (Anthropic sk-ant-\u2026, generic sk-\u2026, Google AIza\u2026, GitHub gh[pousr]_\u2026)\n  { re: /\\bsk-ant-[A-Za-z0-9._\\-]{8,}/g, to: ES_REDACTED },\n  { re: /\\bsk-[A-Za-z0-9._\\-]{16,}/g, to: ES_REDACTED },\n  { re: /\\bAIza[0-9A-Za-z._\\-]{20,}/g, to: ES_REDACTED },\n  { re: /\\bgh[pousr]_[A-Za-z0-9]{20,}/g, to: ES_REDACTED },\n  // JWTs\n  { re: /\\beyJ[A-Za-z0-9._\\-]{10,}\\.[A-Za-z0-9._\\-]{10,}\\.[A-Za-z0-9._\\-]{4,}/g, to: ES_REDACTED },\n  // generic \"<something>key|token|secret|password\" : \"<value>\" (JSON or header style)\n  { re: /(\"?\\b[\\w.\\-]*(?:api[_-]?key|access[_-]?token|refresh[_-]?token|token|secret|password|passwd|pwd|credential)\\b\"?)\\s*[:=]\\s*\"?[^\"\\s,}{&]+\"?/gi, to: '$1: ' + ES_REDACTED },\n  // url query secrets: ?key=\u2026 &token=\u2026\n  { re: /([?&](?:key|token|api_key|apikey|access_token|secret|password)=)[^&\\s]+/gi, to: '$1' + ES_REDACTED },\n  // private PII \u2014 we analyze public positioning, never harvest contacts\n  { re: /[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\\.[a-zA-Z]{2,}/g, to: '[email]' },\n  { re: /\\+?\\d[\\d\\s().-]{9,}\\d/g, to: '[\u0442\u0435\u043b]' }\n];\n\n// Strip a model's hidden reasoning: it is never persisted, in any shape.\nfunction esStripThinking(s) {\n  return esStr(s)\n    .replace(/<thinking>[\\s\\S]*?<\\/thinking>/gi, ' ')\n    .replace(/<thinking>[\\s\\S]*?<\\/antml:thinking>/gi, ' ')\n    .replace(/\"type\"\\s*:\\s*\"thinking\"[\\s\\S]*?(?=[,}]\\s*\"type\"|$)/gi, ' ')\n    .replace(/\\bthinking\\s*[:=]\\s*\"[^\"]*\"/gi, ' ');\n}\n\n// sanitizeErrorPreview(raw, opts) -> a bounded, secret-free, PII-free excerpt safe to persist.\nfunction sanitizeErrorPreview(raw, opts) {\n  opts = opts || {};\n  var max = Number(opts.max_chars);\n  if (!isFinite(max) || max <= 0) max = ES_MAX_PREVIEW;\n  var s = esStripThinking(raw);\n  for (var i = 0; i < ES_RULES.length; i++) s = s.replace(ES_RULES[i].re, ES_RULES[i].to);\n  s = s.replace(/\\s+/g, ' ').trim();\n  if (s.length > max) s = s.slice(0, max - 1) + '\u2026';\n  return s;\n}\n\n// Does a string still look like it carries a secret? Used as a fail-closed assertion before persisting.\nfunction esLooksSecret(s) {\n  s = esStr(s);\n  return /\\b(sk-ant-|sk-[A-Za-z0-9]{16,}|AIza[0-9A-Za-z]{20,}|gh[pousr]_[A-Za-z0-9]{20,}|eyJ[A-Za-z0-9._-]{10,}\\.)/.test(s) ||\n    /(authorization|set-cookie)\\s*[:=]\\s*(?!\\[\u0441\u043a\u0440\u044b\u0442\u043e\\])\\S+/i.test(s);\n}\n\n// The ONLY diagnostic fields a failed record may persist. Anything not listed here is dropped by construction.\nfunction sanitizeErrorRecord(rec, opts) {\n  rec = rec || {};\n  return {\n    provider: esStr(rec.provider),\n    source_url: esStr(rec.source_url),\n    error_category: esStr(rec.error_category || rec.processing_status),\n    parse_error: sanitizeErrorPreview(rec.parse_error, { max_chars: (opts && opts.max_chars) || ES_MAX_PREVIEW }),\n    raw_response_preview: sanitizeErrorPreview(rec.raw_response_preview, { max_chars: (opts && opts.max_chars) || ES_MAX_PREVIEW }),\n    agent_request_id: esStr(rec.agent_request_id),\n    source_run_id: esStr(rec.source_run_id),\n    run_id: esStr(rec.run_id),\n    created_at: esStr(rec.created_at)\n  };\n}\n// --- end embedded error_sanitizer ---\n\nconst data = $json;\nconst srcRecord = $('Normalize Firecrawl Output').first().json;\n// --- Stage C Closure Patch 3: per-record outcome accounting (single point; mutually exclusive; no double count) ---\nconst __sd=$getWorkflowStaticData('global');const __r=(__sd.wf04_run=__sd.wf04_run||{});\n(function(){var pm=String(data.parse_method||'').toLowerCase();var ps=String(data.processing_status||'').toLowerCase();var ru=(data.repair_used===true);\n  if(pm==='deterministic_competitor_fallback'){__r.deterministic_fallback=(__r.deterministic_fallback||0)+1;}\n  else if(pm==='repaired_json'){__r.repair_success=(__r.repair_success||0)+1;}\n  else if(ru && (ps==='technical_error'||String(data.repair_status||'').toLowerCase()==='failed')){__r.repair_failure=(__r.repair_failure||0)+1;}\n})();\n\nfunction clamp(v) { const n = parseInt(v); return isNaN(n) ? 1 : Math.min(100, Math.max(1, n)); }\nfunction truncate(s, n) { return (s == null ? '' : String(s)).substring(0, n); }\n\n// Contact sanitation (DEC-063): partial/placeholder contacts must be blanked. Never invent or keep a partial contact.\n// Empty unless the value matches at least one reliable public-contact pattern (phone +7/8/7 with 10-11 digits, email, telegram, or a contact/profile URL).\n// Deterministic public-contact extraction + sanitation (DEC-070). Keep VALID full contacts; blank\n// partial/hallucinated ones. Prefer a deterministically extracted real contact over a model partial.\n// Reliable types: RU phone (10-11 digits after cleanup; +7/7/8/8-800), email, Telegram (@handle or\n// t.me/), WhatsApp (wa.me/), and contact/profile/application URLs (NOT just the page source_url).\nfunction _normPhoneKey(s){ return String(s||'').replace(/\\D/g,''); }\nfunction extractContacts(textRaw, sourceUrl){\n  var s = (textRaw==null?'':String(textRaw));\n  if (s.trim()==='') return [];\n  var src = (sourceUrl==null?'':String(sourceUrl)).trim().toLowerCase().replace(/\\/+$/,'');\n  var parts = []; var seen = {};\n  // Strip http(s) URLs before scanning phones/emails/bare-handles so digits inside wa.me/<digits>\n  // or t.me/<handle> are not misread as phone numbers or duplicate handles.\n  var sNoUrls = s.replace(/https?:\\/\\/[^\\s,;\"'<>()]+/ig, ' ');\n  // RU phones: +7 / 7 / 8 (incl. 8-800) then 10-11 digits with common separators.\n  var phoneRe = /(?:\\+7|\\b7|\\b8)[\\s()\\-]*\\d[\\d\\s()\\-]{7,}\\d/g; var pm;\n  while ((pm = phoneRe.exec(sNoUrls)) !== null) {\n    var dg = _normPhoneKey(pm[0]);\n    if (dg.length>=10 && dg.length<=11 && !seen['p'+dg]) { seen['p'+dg]=1; parts.push(pm[0].trim()); }\n    if (parts.length>=6) break;\n  }\n  // Emails.\n  var emRe = /[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\\.[A-Za-z]{2,}/g; var em;\n  while ((em = emRe.exec(sNoUrls)) !== null) { var ek='e'+em[0].toLowerCase(); if (!seen[ek]) { seen[ek]=1; parts.push(em[0]); } }\n  // Telegram t.me URLs.\n  var tgRe = /https?:\\/\\/t\\.me\\/[A-Za-z0-9_]+/ig; var tg;\n  while ((tg = tgRe.exec(s)) !== null) { var tk=tg[0].toLowerCase(); if (!seen[tk]) { seen[tk]=1; parts.push(tg[0]); } }\n  // WhatsApp wa.me URLs.\n  var waRe = /https?:\\/\\/wa\\.me\\/\\d+/ig; var wa;\n  while ((wa = waRe.exec(s)) !== null) { var wk=wa[0].toLowerCase(); if (!seen[wk]) { seen[wk]=1; parts.push(wa[0]); } }\n  // Bare Telegram @handle (>=4 chars), not part of an email.\n  var hRe = /(^|[^A-Za-z0-9_@\\/.])@([A-Za-z0-9_]{4,})/g; var h;\n  while ((h = hRe.exec(sNoUrls)) !== null) { var hh='@'+h[2]; var hk=hh.toLowerCase(); if (!seen[hk]) { seen[hk]=1; parts.push(hh); } }\n  // Contact / profile / application URLs (NOT just the source_url).\n  var urlRe = /https?:\\/\\/[^\\s,;\"'<>()]+/ig; var u;\n  while ((u = urlRe.exec(s)) !== null) {\n    var low = u[0].toLowerCase().replace(/[).,;]+$/,'');\n    if (/t\\.me\\/|wa\\.me\\//.test(low)) continue; // already captured above\n    if (!/(api\\.whatsapp\\.com|whatsapp|viber|vk\\.com\\/|contact|kontakt|profile|\\/lk(\\/|$)|zayav|zayavka|order|apply|callback|obratn)/.test(low)) continue;\n    var norm = low.replace(/\\/+$/,'');\n    if (src && norm===src) continue; // not just the page source_url\n    if (!seen['u'+norm]) { seen['u'+norm]=1; parts.push(u[0].replace(/[).,;]+$/,'')); }\n  }\n  return parts.slice(0,6);\n}\n// Best available public contact. Deterministically extract from the MODEL value first (this keeps a\n// valid full model contact and drops partials/ellipsis/\"\u0443\u043a\u0430\u0437\u0430\u043d \u043d\u0430 \u0441\u0430\u0439\u0442\u0435\"/\"\u0442\u0440\u0435\u0431\u0443\u0435\u0442\u0441\u044f \u0438\u0437\u0432\u043b\u0435\u0447\u0435\u043d\u0438\u0435\"); if\n// the model yields nothing reliable, deterministically extract from the page text. Else ''.\nfunction bestContact(modelRaw, textContext, sourceUrl){\n  var fromModel = extractContacts(modelRaw, sourceUrl);\n  if (fromModel.length) return fromModel.join(', ');\n  var fromText = extractContacts(textContext, sourceUrl);\n  return fromText.join(', ');\n}\n\n// Normalize free-text service_type into allowed enum values\nfunction normalizeServiceType(raw, hay) {\n  const allowed = ['secured_auto_loan','secured_real_estate_loan','pts_loan','refinancing','mortgage_adjacent','generic_lending'];\n  const r = (raw || '').toLowerCase().trim();\n  if (allowed.includes(r)) return r;\n  const s = r + ' ' + (hay || '');\n  if (s.includes('\u043f\u0442\u0441') || s.includes('pts')) return 'pts_loan';\n  if ((s.includes('\u0430\u0432\u0442\u043e') || s.includes('\u043c\u0430\u0448\u0438\u043d')) && (s.includes('\u0437\u0430\u043b\u043e\u0433') || s.includes('collateral') || s.includes('\u043b\u043e\u043c\u0431\u0430\u0440\u0434'))) return 'secured_auto_loan';\n  if (s.includes('\u043d\u0435\u0434\u0432\u0438\u0436') || s.includes('\u043a\u0432\u0430\u0440\u0442\u0438\u0440') || s.includes('\u0434\u043e\u043c') || s.includes('\u0437\u0435\u043c\u043b')) return 'secured_real_estate_loan';\n  if (s.includes('\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441') || s.includes('refinanc')) return 'refinancing';\n  if (s.includes('\u0438\u043f\u043e\u0442\u0435\u043a') || s.includes('mortgage')) return 'mortgage_adjacent';\n  if (s.includes('\u0431\u0438\u0437\u043d\u0435\u0441') || s.includes('business')) {\n    if (s.includes('\u043d\u0435\u0434\u0432\u0438\u0436') || s.includes('\u043a\u0432\u0430\u0440\u0442\u0438\u0440')) return 'secured_real_estate_loan';\n    return 'generic_lending';\n  }\n  return 'unknown';\n}\n\n// Descriptive company_name fallback for competitors (never invent a brand)\nfunction companyNameFallback(existing, entity, hay) {\n  if (existing && String(existing).trim() !== '') return existing;\n  if (entity === 'competitor' && (hay.includes('\u043c\u0444\u043e') || hay.includes('\u043c\u0438\u043a\u0440\u043e\u0444\u0438\u043d\u0430\u043d\u0441') || hay.includes('mfo'))) return '\u041c\u0424\u041e / \u0447\u0430\u0441\u0442\u043d\u044b\u0439 \u043a\u0440\u0435\u0434\u0438\u0442\u043e\u0440';\n  if (hay.includes('\u0447\u0430\u0441\u0442\u043d\u044b\u0439 \u0438\u043d\u0432\u0435\u0441\u0442\u043e\u0440')) return '\u0427\u0430\u0441\u0442\u043d\u044b\u0439 \u0438\u043d\u0432\u0435\u0441\u0442\u043e\u0440';\n  if (hay.includes('\u0430\u0432\u0442\u043e\u043b\u043e\u043c\u0431\u0430\u0440\u0434')) return '\u0410\u0432\u0442\u043e\u043b\u043e\u043c\u0431\u0430\u0440\u0434';\n  if (hay.includes('\u0431\u0440\u043e\u043a\u0435\u0440')) return '\u0411\u0440\u043e\u043a\u0435\u0440';\n  if (entity === 'competitor') return '\u041a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442 \u0431\u0435\u0437 \u0431\u0440\u0435\u043d\u0434\u0430';\n  return '';\n}\n\n// recommended_action normalization driven by final route\nfunction normalizeAction(route, entity, action, leadScore) {\n  if (route === 'technical_errors') return 'ignore';\n  if (route === 'skipped_log') return 'ignore';\n  if (route === 'results' && entity === 'lead_signal') return 'contact';\n  if (route === 'review_queue') {\n    if (entity === 'lead_signal' && action === 'contact' && leadScore >= 70) return 'contact';\n    return 'investigate';\n  }\n  if (route === 'monitor_queue' && entity === 'competitor') return 'monitor';\n  if (route === 'content_queue') {\n    if (action === 'contact' || action === 'investigate') return action;\n    return 'create_content';\n  }\n  return action;\n}\n\n// Pass-through for confirmed technical_error (from Parse Repaired JSON)\nif (data.processing_status === 'technical_error' && data.route === 'technical_errors') {\n  return [{ json: {\n    created_at: data.created_at || srcRecord.parsed_at || '',\n    source_type: data.source_type || srcRecord.source_type || '',\n    platform: data.platform || srcRecord.platform || '',\n    source_url: data.source_url || srcRecord.source_url || '',\n    parsed_at: data.parsed_at || srcRecord.parsed_at || '',\n    published_at: data.published_at || srcRecord.published_at || '',\n    freshness_status: data.freshness_status || 'unknown',\n    entity_type: data.entity_type || 'irrelevant',\n    company_name: data.company_name || '',\n    profile_name: data.profile_name || '',\n    profile_url: data.profile_url || srcRecord.profile_url || '',\n    region: data.region || '',\n    service_type: data.service_type || 'unknown',\n    offer_text: data.offer_text || '',\n    terms: data.terms || '',\n    contact_public: bestContact(data.contact_public, (data.text_context || srcRecord.text_context || ''), (data.source_url || srcRecord.source_url || '')),\n    text_context: data.text_context || srcRecord.text_context || '',\n    detected_need: data.detected_need || '',\n    competitor_strength: clamp(data.competitor_strength),\n    lead_signal_score: clamp(data.lead_signal_score),\n    content_idea_score: clamp(data.content_idea_score),\n    quality_score: clamp(data.quality_score),\n    reason: data.reason || '',\n    recommended_action: 'ignore',\n    status: data.status || 'analyzed',\n    processing_status: 'technical_error',\n    parse_method: data.parse_method || 'technical_error',\n    parse_error: sanitizeErrorPreview(data.parse_error || ''),\n    raw_response_preview: sanitizeErrorPreview(data.raw_response_preview),\n    route: 'technical_errors',\n    needs_manual_review: true,\n    repair_used: data.repair_used || false,\n    repair_status: data.repair_status || '',\n    run_id: srcRecord.run_id || '',\n    batch_index: srcRecord.batch_index || 0\n  }}];\n}\n\n// Pass-through for deterministic competitor fallback (DEC-052) \u2014 keep monitor_queue route\nif (data.parse_method === 'deterministic_competitor_fallback') {\n  return [{ json: {\n    created_at: data.created_at || srcRecord.parsed_at || '',\n    source_type: data.source_type || 'scraped_web',\n    platform: data.platform || 'website',\n    source_url: data.source_url || srcRecord.source_url || '',\n    parsed_at: data.parsed_at || srcRecord.parsed_at || '',\n    published_at: data.published_at || '',\n    freshness_status: data.freshness_status || 'unknown',\n    entity_type: 'competitor',\n    company_name: data.company_name || '\u041a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442 \u0431\u0435\u0437 \u0431\u0440\u0435\u043d\u0434\u0430',\n    profile_name: '',\n    profile_url: '',\n    region: data.region || '',\n    service_type: data.service_type || 'generic_lending',\n    offer_text: data.offer_text || '',\n    terms: data.terms || '',\n    contact_public: bestContact(data.contact_public, (data.text_context || srcRecord.text_context || ''), (data.source_url || srcRecord.source_url || '')),\n    text_context: data.text_context || srcRecord.text_context || '',\n    detected_need: '',\n    competitor_strength: clamp(data.competitor_strength),\n    lead_signal_score: clamp(data.lead_signal_score),\n    content_idea_score: clamp(data.content_idea_score),\n    quality_score: clamp(data.quality_score),\n    reason: data.reason || '',\n    recommended_action: 'monitor',\n    status: 'analyzed',\n    processing_status: 'parsed_success',\n    parse_method: 'deterministic_competitor_fallback',\n    parse_error: sanitizeErrorPreview(data.parse_error || ''),\n    raw_response_preview: sanitizeErrorPreview(data.raw_response_preview),\n    route: 'monitor_queue',\n    needs_manual_review: true,\n    repair_used: true,\n    repair_status: data.repair_status || 'failed_fallback',\n    run_id: srcRecord.run_id || data.run_id || '',\n    batch_index: srcRecord.batch_index || data.batch_index || 0\n  }}];\n}\n\n// Normalize schema fields\nconst entity = data.entity_type || 'irrelevant';\nconst status = data.status || 'analyzed';\nconst recAction = data.recommended_action || 'ignore';\nconst leadScore = clamp(data.lead_signal_score);\nconst compScore = clamp(data.competitor_strength);\nconst contentScore = clamp(data.content_idea_score);\nconst qualScore = clamp(data.quality_score);\n\nconst validEntities = ['competitor','lead_signal','market_signal','content_idea','irrelevant'];\nconst validActions = ['monitor','contact','create_content','ignore','investigate'];\nconst safeEntity = validEntities.includes(entity) ? entity : 'irrelevant';\nconst safeAction = validActions.includes(recAction) ? recAction : 'ignore';\nconst safeStatus = (status === 'skipped') ? 'skipped' : 'analyzed';\n\n// Derived text haystacks for normalization and routing keyword checks\nconst hay = ((data.text_context||'') + ' ' + (data.offer_text||'') + ' ' + (data.detected_need||'') + ' ' + (data.reason||'') + ' ' + (data.service_type||'') + ' ' + (srcRecord.text_context||'')).toLowerCase();\nconst companyHay = ((data.text_context||'') + ' ' + (data.reason||'') + ' ' + (data.offer_text||'') + ' ' + (data.source_url||srcRecord.source_url||'') + ' ' + (data.detected_need||'')).toLowerCase();\n\nconst safeServiceType = normalizeServiceType(data.service_type, hay);\nconst safeCompanyName = companyNameFallback(data.company_name, safeEntity, companyHay);\n\n// Determine processing_status\nlet procStatus = 'parsed_success';\nif (safeStatus === 'skipped' || safeEntity === 'irrelevant') procStatus = 'business_skip';\n\n// === Post-repair business-consistency hardening (DEC-043 / DEC-044) ===\n// Repaired JSON is structurally valid but not trusted for business scores/language.\nfunction _hasCyrillic(s){ return /[\\u0400-\\u04FF]/.test(s || ''); }\nfunction _hasCJK(s){ return /[\\u3400-\\u9FFF\\uF900-\\uFAFF]/.test(s || ''); }\n\nconst srcTypeLc = (data.source_type || srcRecord.source_type || '').toLowerCase();\nconst platformLc = (data.platform || srcRecord.platform || '').toLowerCase();\nconst isWebsiteScrape = (srcTypeLc === 'scraped_web') && (platformLc === 'website');\nconst evidence = ((data.text_context||'') + ' ' + (data.offer_text||'') + ' ' + (data.terms||'') + ' ' + (data.reason||'') + ' ' + (srcRecord.text_context||'')).toLowerCase();\nconst hasParseError = !!(data.parse_error && String(data.parse_error).trim() !== '');\nconst textUsable = ((srcRecord.text_context || data.text_context || '').trim().length >= 80);\n\n// Competitor signal counting (B)\nconst signalTerms = ['\u043a\u0440\u0435\u0434\u0438\u0442','\u0437\u0430\u0439\u043c','\u0437\u0430\u043b\u043e\u0433','\u043f\u0442\u0441','\u0430\u0432\u0442\u043e','\u043d\u0435\u0434\u0432\u0438\u0436','\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441','\u0438\u043f\u043e\u0442\u0435\u043a','\u0441\u0442\u0430\u0432\u043a\u0430','\u0441\u0443\u043c\u043c\u0430','\u043e\u0434\u043e\u0431\u0440\u0435\u043d','\u043f\u043b\u043e\u0445\u0430\u044f \u043a\u0440\u0435\u0434\u0438\u0442\u043d','\u043f\u0440\u043e\u0441\u0440\u043e\u0447\u043a','\u0442\u0435\u043b\u0435\u0444\u043e\u043d','\u043c\u043e\u0441\u043a\u0432\u0430','\u043c\u043e\u0441\u043a\u043e\u0432\u0441\u043a'];\nlet compSignals = 0;\nfor (const t of signalTerms) { if (evidence.includes(t)) compSignals++; }\n\n// Effective (hardened) values default to the normalized values\nlet effAction = safeAction;\nlet effComp = compScore;\nlet effQual = qualScore;\nlet effServiceType = safeServiceType;\n\nconst isCompetitorWebsite = (safeEntity === 'competitor' && isWebsiteScrape);\nconst richCompetitor = isCompetitorWebsite && compSignals >= 3;\n\n// B + C: competitor consistency rule (applies even when repaired output under-scored)\nif (isCompetitorWebsite && compSignals >= 3) {\n  effComp = Math.max(effComp, 65);\n  effQual = Math.max(effQual, 65);\n  effAction = 'monitor';\n}\nif (isCompetitorWebsite && compSignals >= 5) {\n  effComp = Math.max(effComp, 75);\n  effQual = Math.max(effQual, 75);\n}\n// C: repaired competitor scores are unreliable \u2014 never below 45 unless text unusable\nif (data.repair_used === true && safeEntity === 'competitor' && textUsable) {\n  effComp = Math.max(effComp, 45);\n}\n\n// D: multi-product website \u2192 generic_lending unless one product dominates\nconst prodCats = [];\nif (evidence.includes('\u0437\u0430\u043b\u043e\u0433') && (evidence.includes('\u043d\u0435\u0434\u0432\u0438\u0436')||evidence.includes('\u043a\u0432\u0430\u0440\u0442\u0438\u0440')||evidence.includes('\u0434\u043e\u043c ')||evidence.includes('\u0437\u0435\u043c\u043b')||evidence.includes('\u043a\u043e\u043c\u043c\u0435\u0440\u0447\u0435\u0441\u043a'))) prodCats.push('real_estate');\nif (evidence.includes('\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441')) prodCats.push('refinancing');\nif (evidence.includes('\u0438\u043f\u043e\u0442\u0435\u043a')) prodCats.push('mortgage');\nif ((evidence.includes('\u0430\u0432\u0442\u043e')||evidence.includes('\u043c\u0430\u0448\u0438\u043d')) && (evidence.includes('\u0437\u0430\u043b\u043e\u0433')||evidence.includes('\u043f\u0442\u0441')||evidence.includes('pts'))) prodCats.push('auto');\nif (isWebsiteScrape && prodCats.length >= 2) {\n  effServiceType = 'generic_lending';\n}\n\n// B2: service_type override (DEC-053/054/062). A specific service page beats the multi-product\n// generic_lending default; a multi-product root homepage stays generic_lending; BUT a root page whose\n// content is overwhelmingly PTS/auto-focused (and not multi-product) becomes pts_loan (DEC-062).\n// svcHay = normalized URL (with path) + evidence (text_context+offer_text+terms+reason, already lowercased).\nconst urlLc = (data.source_url || srcRecord.source_url || '').toLowerCase();\nconst svcHay = urlLc + ' ' + evidence;\nlet urlPath = '';\ntry { urlPath = new URL(urlLc).pathname || '/'; } catch (e) { urlPath = ''; }\nconst isHomepage = (urlPath === '' || urlPath === '/');\n\n// Explicit path/text collateral markers.\nconst ptsTokens = ['pledge-pts','zalog-pts','\u0437\u0430\u043b\u043e\u0433 \u043f\u0442\u0441','\u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433 \u043f\u0442\u0441','\u043f\u0442\u0441 \u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b','\u0437\u0430\u0439\u043c \u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433 \u043f\u0442\u0441'];\nconst ptsExplicit = ptsTokens.some(t => svcHay.includes(t)) || svcHay.includes('\u043f\u0442\u0441') || svcHay.includes('pts');\nconst autoTokens = ['pod-zalog-avto','zalog-avto','\u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e','\u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b','\u043a\u0440\u0435\u0434\u0438\u0442 \u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b'];\nconst autoCollateral = autoTokens.some(t => svcHay.includes(t));\nconst reTokens = ['pod-zalog-nedvizhimosti','zalog-nedvizhimosti','\u0437\u0430\u043b\u043e\u0433 \u043d\u0435\u0434\u0432\u0438\u0436\u0438\u043c\u043e\u0441\u0442','\u043a\u0440\u0435\u0434\u0438\u0442 \u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433 \u043d\u0435\u0434\u0432\u0438\u0436\u0438\u043c\u043e\u0441\u0442','\u043a\u043e\u043c\u043c\u0435\u0440\u0447\u0435\u0441\u043a','\u043a\u0432\u0430\u0440\u0442\u0438\u0440','\u0437\u0430\u043b\u043e\u0433 \u0434\u043e\u043c'];\nconst realEstateCollateral = reTokens.some(t => svcHay.includes(t));\n\n// Deterministic signal counts (distinct tokens present) for content-based override (DEC-062).\nfunction countSig(tokens){ let n = 0; for (const t of tokens) { if (svcHay.includes(t)) n++; } return n; }\nconst ptsAutoSig = countSig(['\u043f\u0442\u0441','\u0437\u0430\u043b\u043e\u0433 \u043f\u0442\u0441','\u043f\u043e\u0434 \u043f\u0442\u0441','\u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e','\u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b','\u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b\u044c \u043e\u0441\u0442\u0430\u0451\u0442\u0441\u044f','\u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b\u044c \u043e\u0441\u0442\u0430\u0435\u0442\u0441\u044f','\u0430\u0432\u0442\u043e\u043b\u043e\u043c\u0431\u0430\u0440\u0434','\u043c\u0430\u0448\u0438\u043d',' \u0430\u0432\u0442\u043e','pts','avto','pledge-pts','pod-zalog-avto','zalog-avto']);\nconst realEstateSig = countSig(['\u043d\u0435\u0434\u0432\u0438\u0436\u0438\u043c\u043e\u0441\u0442','\u043a\u0432\u0430\u0440\u0442\u0438\u0440','\u0434\u043e\u043c','\u0437\u0435\u043c\u043b','\u043a\u043e\u043c\u043c\u0435\u0440\u0447\u0435\u0441\u043a','\u0430\u043f\u0430\u0440\u0442\u0430\u043c\u0435\u043d\u0442','nedvizhimost','pod-zalog-nedvizhimosti']);\nconst refiSig = countSig(['\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441\u0438\u0440\u043e\u0432\u0430\u043d','\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441','refinanc']);\nconst pathAutoPts = /pts|avto|car|zalog/.test(urlLc);\nconst multiProductRoot = (realEstateSig >= 2 && refiSig >= 1) || (ptsAutoSig >= 1 && realEstateSig >= 2 && refiSig >= 1);\n\nif (!isHomepage && ptsExplicit) {\n  effServiceType = 'pts_loan';\n} else if (!isHomepage && autoCollateral) {\n  effServiceType = 'secured_auto_loan';\n} else if (!isHomepage && realEstateCollateral) {\n  effServiceType = 'secured_real_estate_loan';\n} else if (isHomepage && multiProductRoot) {\n  // genuine multi-product root (real estate + refinancing + auto categories) \u2192 keep generic\n  effServiceType = 'generic_lending';\n} else if (ptsAutoSig >= 3 && realEstateSig <= 1 && refiSig === 0) {\n  // overwhelmingly PTS/auto-focused content (applies even to a root homepage) \u2014 DEC-062\n  effServiceType = 'pts_loan';\n} else if (ptsAutoSig >= 3 && pathAutoPts) {\n  effServiceType = 'pts_loan';\n} else if (realEstateSig >= 3 && ptsAutoSig <= 1 && refiSig === 0) {\n  // clearly real-estate-focused content\n  effServiceType = 'secured_real_estate_loan';\n}\n\n// B3: STRONGER final PTS competitor override (this patch). A competitor page whose evidence/URL\n// clearly describes PTS / auto-pawn lending is forced to pts_loan when >=3 distinct strong tokens\n// are present \u2014 UNLESS it is a genuine multi-product root (kept generic_lending) or clearly\n// real-estate-only (kept secured_real_estate_loan). Fixes autolombardn1.ru and\n// autolombard-moskva.ru/services/... while preserving mosinvestfinans root and lioncredit RE page.\nconst ptsStrongTokens = ['\u043f\u0442\u0441','\u044d\u043f\u0442\u0441','\u043f\u043e\u0434 \u043f\u0442\u0441','\u0437\u0430\u043b\u043e\u0433 \u043f\u0442\u0441','\u0437\u0430\u0439\u043c \u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433 \u043f\u0442\u0441','\u0430\u0432\u0442\u043e\u043b\u043e\u043c\u0431\u0430\u0440\u0434','\u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e','\u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b','\u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b\u044f','\u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b\u044c \u043e\u0441\u0442\u0430\u0451\u0442\u0441\u044f','\u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b\u044c \u043e\u0441\u0442\u0430\u0435\u0442\u0441\u044f','\u043c\u0430\u0448\u0438\u043d\u0430 \u043e\u0441\u0442\u0430\u0451\u0442\u0441\u044f','\u043c\u0430\u0448\u0438\u043d\u0430 \u043e\u0441\u0442\u0430\u0435\u0442\u0441\u044f','\u0430\u0432\u0442\u043e \u043e\u0441\u0442\u0430\u0451\u0442\u0441\u044f','\u0430\u0432\u0442\u043e \u043e\u0441\u0442\u0430\u0435\u0442\u0441\u044f','\u0431\u0435\u0437 \u043f\u0440\u043e\u0432\u0435\u0440\u043a\u0438 \u043a\u0438','\u0431\u0435\u0437 \u043f\u0440\u043e\u0432\u0435\u0440\u043a\u0438 \u043a\u0440\u0435\u0434\u0438\u0442\u043d\u043e\u0439 \u0438\u0441\u0442\u043e\u0440\u0438\u0438','\u043b\u044e\u0431\u0430\u044f \u043a\u0438','\u043b\u044e\u0431\u0430\u044f \u043a\u0440\u0435\u0434\u0438\u0442\u043d\u0430\u044f \u0438\u0441\u0442\u043e\u0440\u0438\u044f'];\nlet ptsStrongHits = 0;\nfor (const t of ptsStrongTokens) { if (svcHay.includes(t)) ptsStrongHits++; }\nconst realEstateOnly = (realEstateSig >= 3 && ptsAutoSig <= 1) || (realEstateCollateral && !ptsExplicit && ptsAutoSig === 0);\nif (safeEntity === 'competitor' && ptsStrongHits >= 3 && !multiProductRoot && !realEstateOnly) {\n  effServiceType = 'pts_loan';\n}\n\n// A: language guard \u2014 Russian source but reason in Chinese / mostly foreign script\nfunction _russianFallbackReason(entity, company, service, terms, region, lead, comp, qual) {\n  if (entity === 'competitor') {\n    return '\u0421\u0442\u0440\u0430\u043d\u0438\u0446\u0430 \u0441\u043e\u0434\u0435\u0440\u0436\u0438\u0442 \u043f\u0440\u0438\u0437\u043d\u0430\u043a\u0438 \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u0430 \u0432 \u043d\u0438\u0448\u0435 \u0437\u0430\u043b\u043e\u0433\u043e\u0432\u043e\u0433\u043e \u043a\u0440\u0435\u0434\u0438\u0442\u043e\u0432\u0430\u043d\u0438\u044f: \u0443\u043a\u0430\u0437\u0430\u043d\u044b \u043f\u0440\u043e\u0434\u0443\u043a\u0442\u044b, \u0443\u0441\u043b\u043e\u0432\u0438\u044f, \u0441\u0443\u043c\u043c\u044b/\u0441\u0442\u0430\u0432\u043a\u0438 \u0438\u043b\u0438 \u043a\u043e\u043d\u0442\u0430\u043a\u0442. \u0417\u0430\u043f\u0438\u0441\u044c \u043e\u0442\u043f\u0440\u0430\u0432\u043b\u0435\u043d\u0430 \u0432 \u043c\u043e\u043d\u0438\u0442\u043e\u0440\u0438\u043d\u0433 \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u043e\u0432.';\n  }\n  const parts = ['\u0410\u0432\u0442\u043e\u043c\u0430\u0442\u0438\u0447\u0435\u0441\u043a\u0438 \u0441\u0444\u043e\u0440\u043c\u0438\u0440\u043e\u0432\u0430\u043d\u043d\u043e\u0435 \u043e\u043f\u0438\u0441\u0430\u043d\u0438\u0435 (\u0438\u0441\u0445\u043e\u0434\u043d\u044b\u0439 reason \u0431\u044b\u043b \u043d\u0435 \u043d\u0430 \u0440\u0443\u0441\u0441\u043a\u043e\u043c).'];\n  parts.push('\u0422\u0438\u043f \u0437\u0430\u043f\u0438\u0441\u0438: ' + entity + (company ? ', ' + company : '') + (service && service !== 'unknown' ? ', \u0443\u0441\u043b\u0443\u0433\u0430: ' + service : '') + '.');\n  if (region) parts.push('\u0420\u0435\u0433\u0438\u043e\u043d: ' + region + '.');\n  if (terms) parts.push('\u0423\u0441\u043b\u043e\u0432\u0438\u044f: ' + truncate(terms, 120) + '.');\n  parts.push('\u041e\u0446\u0435\u043d\u043a\u0438 \u2014 \u043b\u0438\u0434: ' + lead + ', \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442: ' + comp + ', \u043a\u0430\u0447\u0435\u0441\u0442\u0432\u043e: ' + qual + '.');\n  return parts.join(' ');\n}\nlet finalReason = data.reason || '';\nconst srcHasCyrillic = _hasCyrillic(srcRecord.text_context || data.text_context || evidence);\nconst reasonLetters = (finalReason.match(/[A-Za-z\\u0400-\\u04FF]/g) || []).length;\nconst reasonNonSpace = (finalReason.match(/\\S/g) || []).length;\nconst reasonForeign = reasonNonSpace >= 8 && (reasonLetters / reasonNonSpace) < 0.4;\nconst reasonCyr = (finalReason.match(/[\\u0400-\\u04FF]/g) || []).length;\nconst reasonLatin = (finalReason.match(/[A-Za-z]/g) || []).length;\nconst reasonMostlyEnglish = reasonLatin >= 8 && (reasonCyr / Math.max(1, reasonLatin + reasonCyr)) < 0.3;\nif (srcHasCyrillic && finalReason && (_hasCJK(finalReason) || reasonForeign || reasonMostlyEnglish || (!_hasCyrillic(finalReason) && reasonLetters === 0))) {\n  finalReason = _russianFallbackReason(safeEntity, safeCompanyName, effServiceType, data.terms || '', data.region || '', leadScore, effComp, effQual);\n}\n\n// A2: offer_text language guard (DEC-053) \u2014 Russian source but English/foreign offer_text.\nlet finalOffer = data.offer_text || '';\nconst offerLatin = (finalOffer.match(/[A-Za-z]/g) || []).length;\nconst offerCyr = (finalOffer.match(/[\\u0400-\\u04FF]/g) || []).length;\nconst offerMostlyEnglish = offerLatin >= 8 && (offerCyr / Math.max(1, offerLatin + offerCyr)) < 0.3;\nif (srcHasCyrillic && finalOffer && (offerMostlyEnglish || _hasCJK(finalOffer))) {\n  const _svc = (effServiceType && effServiceType !== 'unknown') ? effServiceType : '\u043a\u0440\u0435\u0434\u0438\u0442\u043d\u044b\u0439 \u043f\u0440\u043e\u0434\u0443\u043a\u0442';\n  const _comp = safeCompanyName || '\u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442';\n  const _terms = (data.terms && String(data.terms).trim() !== '') ? (' \u0423\u0441\u043b\u043e\u0432\u0438\u044f: ' + truncate(data.terms, 120) + '.') : '';\n  finalOffer = '\u041f\u0440\u0435\u0434\u043b\u043e\u0436\u0435\u043d\u0438\u0435 \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u0430 (' + _comp + '), \u0443\u0441\u043b\u0443\u0433\u0430: ' + _svc + '.' + _terms;\n}\n\n// Weak / potential lead detection (must run before content_queue routing)\nconst srcType = (data.source_type || srcRecord.source_type || '').toLowerCase();\nconst isSocialOrClassified = (srcType === 'social' || srcType === 'classified');\nconst productTerms = ['loan','collateral','pts','auto','real estate','refinanc','\u0437\u0430\u0439\u043c','\u043a\u0440\u0435\u0434\u0438\u0442','\u0437\u0430\u043b\u043e\u0433','\u043f\u0442\u0441','\u0430\u0432\u0442\u043e','\u043c\u0430\u0448\u0438\u043d','\u043d\u0435\u0434\u0432\u0438\u0436','\u043a\u0432\u0430\u0440\u0442\u0438\u0440','\u0437\u0435\u043c\u043b','\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441','\u0438\u043f\u043e\u0442\u0435\u043a'];\nconst mentionsProduct = productTerms.some(t => hay.includes(t));\n\nconst weakLead = (\n  (safeEntity === 'lead_signal' && leadScore >= 30 && leadScore <= 69) ||\n  (effAction === 'investigate') ||\n  (leadScore >= 30 && isSocialOrClassified && mentionsProduct) ||\n  (safeEntity === 'content_idea' && leadScore >= 30 && isSocialOrClassified && safeServiceType !== 'unknown')\n);\n\n// Routing priority (F, after consistency hardening \u2014 DEC-043):\n// 1 technical_errors (handled above) 2 skipped/irrelevant 3 hot lead\n// 4 rich competitor website (never review_queue, rule C) 5 weak lead\n// 6 competitor>=45 7 pure content idea 8 fallback review_queue\nlet route;\nlet needsReview;\n\nif (safeStatus === 'skipped' || safeEntity === 'irrelevant') {\n  route = 'skipped_log'; needsReview = false;\n} else if (safeEntity === 'lead_signal' && leadScore >= 70 && effAction === 'contact') {\n  route = 'results'; needsReview = false;\n} else if (richCompetitor) {\n  route = 'monitor_queue'; needsReview = hasParseError;\n} else if (weakLead) {\n  route = 'review_queue'; needsReview = true;\n} else if (safeEntity === 'competitor' && effComp >= 45) {\n  route = 'monitor_queue'; needsReview = false;\n} else if (safeEntity === 'content_idea' && contentScore >= 50) {\n  route = 'content_queue'; needsReview = true;\n} else {\n  route = 'review_queue'; needsReview = true;\n}\n\n// Safety: enforce route is exactly one of the six valid sheet tabs\nconst validRoutes = ['results','review_queue','monitor_queue','content_queue','skipped_log','technical_errors'];\nlet parseError = data.parse_error || '';\nif (!validRoutes.includes(route)) {\n  route = 'technical_errors';\n  procStatus = 'technical_error';\n  needsReview = true;\n  parseError = (parseError ? parseError + '; ' : '') + 'invalid_route';\n}\n\n// Final recommended_action consistent with the route\nconst finalAction = normalizeAction(route, safeEntity, effAction, leadScore);\n\n// v0.1 dedup hint: source_url is the first dedup key. Real scraper workflow\n// should check existing source_url in the target tab before append. No dedup\n// column is emitted yet (see DEC-037 / TABLE_SCHEMA notes).\n\nreturn [{ json: {\n  created_at: data.created_at || srcRecord.parsed_at || '',\n  source_type: data.source_type || srcRecord.source_type || '',\n  platform: data.platform || srcRecord.platform || '',\n  source_url: data.source_url || srcRecord.source_url || '',\n  parsed_at: data.parsed_at || srcRecord.parsed_at || '',\n  published_at: data.published_at || srcRecord.published_at || '',\n  freshness_status: data.freshness_status || 'unknown',\n  entity_type: safeEntity,\n  company_name: safeCompanyName,\n  profile_name: data.profile_name || '',\n  profile_url: data.profile_url || srcRecord.profile_url || '',\n  region: data.region || '',\n  service_type: effServiceType,\n  offer_text: finalOffer,\n  terms: data.terms || '',\n  contact_public: bestContact(data.contact_public, (data.text_context || srcRecord.text_context || ''), (data.source_url || srcRecord.source_url || '')),\n  text_context: data.text_context || '',\n  detected_need: data.detected_need || '',\n  competitor_strength: effComp,\n  lead_signal_score: leadScore,\n  content_idea_score: contentScore,\n  quality_score: effQual,\n  reason: finalReason,\n  recommended_action: finalAction,\n  status: safeStatus,\n  processing_status: procStatus,\n  parse_method: data.parse_method || 'primary_json',\n  parse_error: sanitizeErrorPreview(parseError),\n  raw_response_preview: sanitizeErrorPreview(data.raw_response_preview),\n  route: route,\n  needs_manual_review: needsReview,\n  repair_used: data.repair_used || false,\n  repair_status: data.repair_status || '',\n  run_id: srcRecord.run_id || '',\n  batch_index: srcRecord.batch_index || 0\n}}];"
      }
    },
    {
      "id": "b4000000-0000-0000-0000-0000000000sk",
      "name": "Append Skipped Log (Duplicate)",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        1540,
        520
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "={{ $json.route }}",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000,
      "onError": "continueRegularOutput",
      "alwaysOutputData": true
    },
    {
      "id": "rr000000-0000-0000-0000-0000000000rb",
      "name": "Build Registry Row",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        4180,
        300
      ],
      "parameters": {
        "jsCode": "const row = $json;\nconst ctx = $('Normalize URL for Dedup').first().json;\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nconst now = moscowIsoNow();\n// One url_registry row per non-duplicate processing attempt (DEC-051).\nreturn [{ json: {\n  normalized_source_url: ctx.normalized_source_url || row.source_url || '',\n  source_url: row.source_url || ctx.source_url || '',\n  first_seen_at: now,\n  last_seen_at: now,\n  last_route: row.route || '',\n  last_processing_status: row.processing_status || '',\n  last_entity_type: row.entity_type || '',\n  run_id: row.run_id || ctx.run_id || '',\n  batch_index: row.batch_index || ctx.batch_index || 0,\n  note: 'processed_by_workflow_04'\n}}];"
      }
    },
    {
      "id": "rr000000-0000-0000-0000-0000000000ra",
      "name": "Append url_registry",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        4400,
        300
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "url_registry",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "rr000000-0000-0000-0000-000000000099",
      "name": "Append to Dynamic Route Sheet",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        3960,
        300
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "={{ $json.route }}",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000,
      "onError": "continueRegularOutput",
      "alwaysOutputData": true
    },
    {
      "id": "fc000000-0000-0000-0000-0000000000n2",
      "name": "Test Instructions RU",
      "type": "n8n-nodes-base.stickyNote",
      "typeVersion": 1,
      "position": [
        4600,
        -160
      ],
      "parameters": {
        "content": "## \u0422\u0435\u0441\u0442 (\u0432\u0440\u0443\u0447\u043d\u0443\u044e)\n\n1. \u041d\u0415 \u0430\u043a\u0442\u0438\u0432\u0438\u0440\u043e\u0432\u0430\u0442\u044c workflow.\n2. \u0421\u043e\u0437\u0434\u0430\u0442\u044c \u0432\u043a\u043b\u0430\u0434\u043a\u0443 url_registry \u0441 10 \u043a\u043e\u043b\u043e\u043d\u043a\u0430\u043c\u0438 (docs/TABLE_SCHEMA.md): normalized_source_url, source_url, first_seen_at, last_seen_at, last_route, last_processing_status, last_entity_type, run_id, batch_index, note.\n3. \u041f\u041e\u0421\u041b\u0415 \u041a\u0410\u0416\u0414\u041e\u0413\u041e \u0418\u041c\u041f\u041e\u0420\u0422\u0410 \u043f\u0435\u0440\u0435\u043f\u0440\u0438\u0432\u044f\u0437\u0430\u0442\u044c \u043a\u0440\u0435\u0434\u0435\u043d\u0448\u043b\u044b (ID \u043b\u043e\u043a\u0430\u043b\u044c\u043d\u044b):\n   - Firecrawl Scrape API \u2192 Firecrawl API - Marketing Scout\n   - Claude Primary / Claude Repair \u2192 Claude API - Marketing Scout\n   - Registry Lookup / Append url_registry / Append Skipped Log (Duplicate) / Append to Dynamic Route Sheet \u2192 Google Sheets - Marketing Scout Service Account\n   - Append competitor_site_snapshots / Append live_source_runs / Append agent_requests \u2192 Google Sheets - Marketing Scout Service Account\n4. \u041d\u0430 \u0432\u0441\u0435\u0445 Google Sheets \u043d\u043e\u0434\u0430\u0445 \u0432\u0441\u0442\u0430\u0432\u0438\u0442\u044c \u0440\u0435\u0430\u043b\u044c\u043d\u044b\u0439 Spreadsheet ID.\n4b. \u0412\u043a\u043b\u0430\u0434\u043a\u0438 \u043d\u0430\u0431\u043b\u044e\u0434\u0430\u0435\u043c\u043e\u0441\u0442\u0438/\u0441\u043d\u0430\u043f\u0448\u043e\u0442\u043e\u0432 (docs/TABLE_SCHEMA.md): competitor_site_snapshots (22 \u043a\u043e\u043b.), live_source_runs (23 \u043a\u043e\u043b.), agent_requests (21 \u043a\u043e\u043b.) \u2014 \u0434\u043e\u043b\u0436\u043d\u044b \u0441\u0443\u0449\u0435\u0441\u0442\u0432\u043e\u0432\u0430\u0442\u044c. \u0421\u043d\u0430\u043f\u0448\u043e\u0442\u044b \u043f\u0438\u0448\u0443\u0442\u0441\u044f per-URL (baseline, change_type=baseline); live_source_runs/agent_requests \u2014 \u043e\u0434\u043d\u0430 \u0441\u0442\u0440\u043e\u043a\u0430 \u043d\u0430 \u043f\u0440\u043e\u0433\u043e\u043d (\u0432\u0435\u0442\u043a\u0430 done \u0446\u0438\u043a\u043b\u0430).\n5. 6 \u0431\u0438\u0437\u043d\u0435\u0441-\u0432\u043a\u043b\u0430\u0434\u043e\u043a \u0441 35-\u043a\u043e\u043b\u043e\u043d\u043e\u0447\u043d\u044b\u043c \u0437\u0430\u0433\u043e\u043b\u043e\u0432\u043a\u043e\u043c + \u0432\u043a\u043b\u0430\u0434\u043a\u0430 url_registry (10 \u043a\u043e\u043b\u043e\u043d\u043e\u043a).\n6. \u0412 'Set URL List' \u0432\u0441\u0442\u0430\u0432\u0438\u0442\u044c 3 URL (\u043f\u0435\u0440\u0432\u044b\u0439 \u0437\u0430\u043f\u0443\u0441\u043a). \u041c\u0430\u043a\u0441\u0438\u043c\u0443\u043c 5 (\u0436\u0451\u0441\u0442\u043a\u0430\u044f \u043e\u0442\u0441\u0435\u0447\u043a\u0430).\n7. \u0417\u0430\u043f\u0438\u0441\u0430\u0442\u044c \u043a\u0440\u0435\u0434\u0438\u0442\u044b Firecrawl \u0438 \u0431\u0430\u043b\u0430\u043d\u0441 Claude \u0414\u041e.\n8. Execute Workflow \u043e\u0434\u0438\u043d \u0440\u0430\u0437. \u0417\u0430\u043f\u0438\u0441\u0430\u0442\u044c \u041f\u041e\u0421\u041b\u0415.\n\n\u041e\u0436\u0438\u0434\u0430\u0435\u043c\u043e:\n- \u041d\u043e\u0432\u044b\u0439 \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442 \u2192 monitor_queue + \u0441\u0442\u0440\u043e\u043a\u0430 \u0432 url_registry.\n- \u041f\u043e\u0432\u0442\u043e\u0440\u043d\u044b\u0439 URL (\u0435\u0441\u0442\u044c \u0432 url_registry) \u2192 skipped_log, parse_method=dedup_source_url, \u0411\u0415\u0417 \u0442\u0440\u0430\u0442.\n- Firecrawl \u0443\u043f\u0430\u043b/\u043f\u0443\u0441\u0442\u043e \u2192 technical_errors (Claude \u043d\u0435 \u0432\u044b\u0437\u044b\u0432\u0430\u0435\u0442\u0441\u044f), \u0441\u0442\u0440\u043e\u043a\u0430 \u0432 url_registry.\n- \u041f\u0440\u0430\u0439\u043c\u0435\u0440\u0438+\u0440\u0435\u043c\u043e\u043d\u0442 \u0431\u0435\u0437 JSON, \u043d\u043e \u22655 \u0441\u0438\u0433\u043d\u0430\u043b\u043e\u0432 \u2192 deterministic_competitor_fallback \u2192 monitor_queue.\n\n\u0420\u0435\u0442\u0435\u0441\u0442: \u043e\u0441\u0442\u0430\u0432\u0438\u0442\u044c \u0437\u0430\u043f\u0438\u0441\u044c url_registry \u0434\u043b\u044f \u043e\u0434\u043d\u043e\u0433\u043e \u0438\u0437 example.com URL, \u0437\u0430\u043f\u0443\u0441\u0442\u0438\u0442\u044c \u0442\u0435 \u0436\u0435 3 URL \u043f\u043e\u0432\u0442\u043e\u0440\u043d\u043e \u2192 \u043e\u0436\u0438\u0434\u0430\u0442\u044c \u0434\u0443\u0431\u043b\u0438\u043a\u0430\u0442 (skipped_log) \u0434\u043b\u044f \u0432\u0441\u0435\u0445 \u0443\u0436\u0435 \u043e\u0431\u0440\u0430\u0431\u043e\u0442\u0430\u043d\u043d\u044b\u0445, monitor_queue \u0442\u043e\u043b\u044c\u043a\u043e \u0434\u043b\u044f \u043d\u043e\u0432\u044b\u0445. \u0415\u0441\u043b\u0438 lookup/append \u043f\u0430\u0434\u0430\u044e\u0442 \u043f\u0440\u0438 \u0438\u043c\u043f\u043e\u0440\u0442\u0435 \u2014 \u0441\u043c. docs/N8N_WORKFLOW_04_FIRECRAWL_URL_LIST_RU.md (fallback).",
        "height": 520,
        "width": 540,
        "color": 4
      }
    },
    {
      "id": "wf04-build-snap",
      "name": "Build competitor_site_snapshots Row",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1500,
        1040
      ],
      "parameters": {
        "jsCode": "function moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nfunction moscowStamp(){var z=function(n){return String(n).padStart(2,'0');};var m=new Date(Date.now()+10800000);return m.getUTCFullYear()+z(m.getUTCMonth()+1)+z(m.getUTCDate())+'_'+z(m.getUTCHours())+z(m.getUTCMinutes())+z(m.getUTCSeconds());}\n// WF04 -> competitor_site_snapshots writer (Stage C Closure Patch 2). Per analyzed competitor WEBSITE page.\n// Evidence-based: brand is preserved from title/domain when analysis left it blank (NEVER \"\u041a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442 \u0431\u0435\u0437 \u0431\u0440\u0435\u043d\u0434\u0430\"\n// when title/domain evidence exists, S2-D10/D12); page_type uses URL + title/content (S2-D14); canonical\n// service_primary/service_secondary (S2-D13); confidence reflects evidence completeness (S2-D11); phones are\n// normalized to canonical RU form (S2-D15); raw Markdown is kept ONLY in an audit field, never in the\n// stakeholder offer field (S2-D10/D20); quality_status/report_eligible/fallback_reason gate the report;\n// cost telemetry is explicit with cost_status=unknown when not recovered (S2-D17). No external calls.\nfunction str(v){return v==null?'':String(v).trim();}\nfunction low(v){return str(v).toLowerCase();}\nfunction cut(s,n){s=str(s);return s.length>n?s.slice(0,n):s;}\nfunction hash(s){s=str(s).toLowerCase().replace(/\\s+/g,' ');var h=0;for(var i=0;i<s.length;i++){h=((h<<5)-h+s.charCodeAt(i))|0;}return 'h'+(h>>>0).toString(16);}\nfunction sheetSafe(v){v=str(v);return (v && /^[=+\\-@0-9]/.test(v))?(\"'\"+v):v;}\nconst r=$json;const route=low(r.route);const ps=low(r.processing_status);\nif(ps==='technical_error'||route==='technical_errors'||route==='skipped_log') return [];\nconst src=str(r.source_url); if(!src) return [];\nfunction domainOf(u){var h='';try{h=new URL(u).hostname;}catch(e){var m=String(u).match(/^https?:\\/\\/([^\\/?#]+)/i);h=m?m[1]:'';}h=h.toLowerCase();if(h.indexOf('www.')===0)h=h.slice(4);return h;}\nconst domain=domainOf(src);\nconst title=str(r.title||r.page_title);\nconst txt=str(r.text_context);\nconst offerRaw=str(r.offer_text);\nconst blob=(title+' '+cut(txt,600)+' '+cut(offerRaw,400)).toLowerCase();\n\n// ---- raw-markdown guard (S2-D10/D20): never put huge raw markdown into the stakeholder offer field ----\nfunction looksRawMarkdown(s){s=str(s);return /(^|\\n)\\s{0,3}#{1,6}\\s/.test(s)||/\\]\\(https?:/.test(s)||/(^|\\n)\\s*[-*]\\s+\\S/.test(s)||/\\|\\s*-{2,}\\s*\\|/.test(s)||s.length>600;}\nfunction plainText(s){return str(s).replace(/```[\\s\\S]*?```/g,' ').replace(/!\\[[^\\]]*\\]\\([^)]*\\)/g,' ').replace(/\\[([^\\]]*)\\]\\([^)]*\\)/g,'$1').replace(/[#*`>_|]+/g,' ').replace(/https?:\\/\\/\\S+/g,' ').replace(/\\s+/g,' ').trim();}\nconst offerIsRaw=looksRawMarkdown(offerRaw);\nconst fallbackParse=(low(r.parse_method)==='deterministic_competitor_fallback'||low(r.repair_status)==='failed_fallback');\n\n// ---- brand preservation (S2-D10/D12) ----\nfunction brandFromDomain(d){d=str(d).replace(/^www\\./,'').split('.')[0];if(!d)return '';return d.charAt(0).toUpperCase()+d.slice(1);}\nfunction brandFromTitle(t){t=str(t);if(!t)return '';var seg=t.split(/[|\u2013\u2014:\u2022\u00b7\u00bb\\-]/)[0].trim();if(seg.length<2)return '';return cut(seg,60);}\nlet company=cut(str(r.company_name),120);\nlet brand_source='analysis';\nif(!company||/^\u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442 \u0431\u0435\u0437 \u0431\u0440\u0435\u043d\u0434\u0430$/i.test(company)){\n  const bt=brandFromTitle(title), bd=brandFromDomain(domain);\n  if(bt){company=bt;brand_source='title';}\n  else if(bd){company=bd;brand_source='domain';}\n  else {company='\u041a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442 \u0431\u0435\u0437 \u0431\u0440\u0435\u043d\u0434\u0430';brand_source='none';}\n}\nconst broken_brand=(brand_source==='none');\n\n// ---- page_type from URL + title/content evidence (S2-D14), canonical ----\nfunction pageType(u){\n  var p='';try{p=new URL(u).pathname.toLowerCase();}catch(e){p=String(u).toLowerCase();}\n  if(/contact|kontakt|\\/lk(\\/|$)/.test(p)||/\u043a\u043e\u043d\u0442\u0430\u043a\u0442/.test(blob))return 'contact';\n  if(/price|tarif|\\bcena\\b|cost|tarify/.test(p)||/\u0442\u0430\u0440\u0438\u0444|\u0441\u0442\u043e\u0438\u043c\u043e\u0441\u0442|\\b\u0446\u0435\u043d\u044b\\b|\u043f\u0440\u0430\u0439\u0441/.test(blob))return 'prices';\n  if(/about|o-nas|o_nas|company|o_kompanii/.test(p)||/\u043e \u043a\u043e\u043c\u043f\u0430\u043d\u0438\u0438|\u043e \u043d\u0430\u0441/.test(blob))return 'about';\n  if(/service|uslugi|zaim|zaym|kredit|broker|ipotek|refinans|pts/.test(p))return 'services';\n  if(p===''||p==='/')return 'home';\n  if(/\u043e\u0441\u0442\u0430\u0432\u044c\u0442\u0435 \u0437\u0430\u044f\u0432\u043a\u0443|\u043e\u0441\u0442\u0430\u0432\u0438\u0442\u044c \u0437\u0430\u044f\u0432\u043a\u0443|\u043e\u0444\u043e\u0440\u043c\u0438\u0442\u044c \u0437\u0430\u044f\u0432\u043a\u0443|\u043f\u043e\u043b\u0443\u0447\u0438\u0442\u044c \u0434\u0435\u043d\u044c\u0433\u0438|\u043e\u043d\u043b\u0430\u0439\u043d[- ]?\u0437\u0430\u044f\u0432\u043a/.test(blob))return 'offer';\n  if(/\u0443\u0441\u043b\u0443\u0433|\u043a\u0440\u0435\u0434\u0438\u0442|\u0437\u0430\u0439\u043c|\u0431\u0440\u043e\u043a\u0435\u0440|\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441|\u0438\u043f\u043e\u0442\u0435\u043a|\u0437\u0430\u043b\u043e\u0433|\u043f\u0442\u0441/.test(blob))return 'services';\n  return (title||txt)?'home':'other';\n}\nconst page_type=pageType(src);\n\n// ---- canonical service detection (S2-D13). Broad brokerage evidence wins over a narrow product. ----\nconst SVC=[\n  ['credit_brokerage',['\u043a\u0440\u0435\u0434\u0438\u0442\u043d\u044b\u0439 \u0431\u0440\u043e\u043a\u0435\u0440','\u043a\u0440\u0435\u0434\u0438\u0442\u043d\u044b\u0435 \u0431\u0440\u043e\u043a\u0435\u0440','\u043f\u043e\u043c\u043e\u0449\u044c \u0432 \u043f\u043e\u043b\u0443\u0447\u0435\u043d\u0438','\u043f\u043e\u0434\u0431\u043e\u0440 \u0431\u0430\u043d\u043a','\u043f\u043e\u043c\u043e\u0449\u044c \u0441 \u043a\u0440\u0435\u0434\u0438\u0442','\u043e\u0434\u043e\u0431\u0440\u0435\u043d\u0438\u0435 \u043a\u0440\u0435\u0434\u0438\u0442','broker','\u0431\u0440\u043e\u043a\u0435\u0440']],\n  ['credit_after_refusals',['\u043f\u043e\u0441\u043b\u0435 \u043e\u0442\u043a\u0430\u0437\u043e\u0432','\u043e\u0442\u043a\u0430\u0437\u0430\u043b\u0438','\u0441 \u043f\u043b\u043e\u0445\u043e\u0439 \u043a\u0440\u0435\u0434\u0438\u0442\u043d','\u043f\u0440\u043e\u0441\u0440\u043e\u0447\u043a']],\n  ['mortgage_brokerage',['\u0438\u043f\u043e\u0442\u0435\u043a','\u0438\u043f\u043e\u0442\u0435\u0447\u043d']],\n  ['mortgage_refinance',['\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441\u0438\u0440\u043e\u0432\u0430\u043d\u0438\u0435 \u0438\u043f\u043e\u0442\u0435\u043a','\u0441\u043d\u0438\u0436\u0435\u043d\u0438\u0435 \u0441\u0442\u0430\u0432\u043a\u0438 \u043f\u043e \u0438\u043f\u043e\u0442\u0435\u043a']],\n  ['debt_refinancing',['\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441\u0438\u0440','\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441\u0438\u0440\u043e\u0432\u0430\u043d\u0438\u0435 \u043a\u0440\u0435\u0434\u0438\u0442']],\n  ['pts_loan',['\u043f\u043e\u0434 \u043f\u0442\u0441','\u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e','\u0437\u0430\u0439\u043c \u043f\u043e\u0434 \u0430\u0432\u0442\u043e','\u0434\u0435\u043d\u044c\u0433\u0438 \u043f\u043e\u0434 \u043f\u0442\u0441','pts']],\n  ['real_estate_secured_loan',['\u0437\u0430\u043b\u043e\u0433 \u043d\u0435\u0434\u0432\u0438\u0436','\u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433 \u043a\u0432\u0430\u0440\u0442\u0438\u0440','\u043f\u043e\u0434 \u043d\u0435\u0434\u0432\u0438\u0436']],\n  ['business_credit',['\u0434\u043b\u044f \u0431\u0438\u0437\u043d\u0435\u0441\u0430','\u043e\u0431\u043e\u0440\u043e\u0442\u043d','\u0442\u0435\u043d\u0434\u0435\u0440\u043d','\u0431\u0438\u0437\u043d\u0435\u0441-\u043a\u0440\u0435\u0434\u0438\u0442','\u0438\u043f \u0438 \u043e\u043e\u043e']],\n  ['bank_guarantee',['\u0431\u0430\u043d\u043a\u043e\u0432\u0441\u043a \u0433\u0430\u0440\u0430\u043d\u0442','\u0431\u0430\u043d\u043a\u043e\u0432\u0441\u043a\u0438\u0435 \u0433\u0430\u0440\u0430\u043d\u0442\u0438\u0438']],\n  ['auto_credit',['\u0430\u0432\u0442\u043e\u043a\u0440\u0435\u0434\u0438\u0442']],\n  ['consumer_credit',['\u043f\u043e\u0442\u0440\u0435\u0431\u0438\u0442\u0435\u043b\u044c\u0441\u043a \u043a\u0440\u0435\u0434\u0438\u0442','\u043f\u043e\u0442\u0440\u0435\u0431 \u043a\u0440\u0435\u0434\u0438\u0442']]\n];\nconst svcHits=[];for(const pair of SVC){if(pair[1].some(t=>blob.indexOf(t)>=0))svcHits.push(pair[0]);}\nfunction canonService(s){s=low(s);const A={'credit_broker':'credit_brokerage','kreditnyy_broker':'credit_brokerage','generic_lending':'credit_brokerage','secured_auto_loan':'pts_loan','auto_collateral_loan':'pts_loan','refinancing':'debt_refinancing','ipoteka':'mortgage_brokerage','consumer_loan':'consumer_credit'};return A[s]||s;}\n// seed with the analysis-provided service (canonicalized), then evidence hits \u2014 broadest brokerage first\nlet services=[];\nconst provided=canonService(str(r.service_type));\nif(provided&&provided!=='unknown'&&provided!=='generic_lending')services.push(provided);\nfor(const h of svcHits)if(services.indexOf(h)<0)services.push(h);\n// if broad brokerage evidence exists, it must be primary (do NOT narrow Finardi to pts_loan)\nif(services.indexOf('credit_brokerage')>0){services.splice(services.indexOf('credit_brokerage'),1);services.unshift('credit_brokerage');}\nif(!services.length)services=['unknown'];\nconst service_primary=services[0];\nconst service_secondary=services.slice(1,3).join(', ');\n\n// ---- phone normalization (S2-D15): canonical RU; dedupe; never invent ----\nfunction normalizePhonesRu(s){\n  s=str(s);const found=[];const seen={};\n  const re=/(?:\\+7|\\b8|\\b7)[\\s()\\-]*\\d[\\d\\s()\\-]{8,}\\d/g;let m;\n  while((m=re.exec(s))!==null){\n    let d=m[0].replace(/\\D/g,'');\n    if(d.length===11&&(d[0]==='7'||d[0]==='8'))d='7'+d.slice(1);\n    else if(d.length===10)d='7'+d;\n    else continue;\n    if(d.length!==11)continue;\n    if(seen[d])continue;seen[d]=1;\n    found.push('+7 ('+d.slice(1,4)+') '+d.slice(4,7)+'-'+d.slice(7,9)+'-'+d.slice(9,11));\n  }\n  return found;\n}\nconst contactRaw=str(r.contact_public);\nconst phones=normalizePhonesRu(contactRaw);\nlet contactNorm=contactRaw;\nif(phones.length){\n  // replace the phone-ish span(s) with canonical forms; keep any non-phone identifiers as-is\n  contactNorm=contactRaw.replace(/(?:\\+7|\\b8|\\b7)[\\s()\\-]*\\d[\\d\\s()\\-]{8,}\\d/g,' ').replace(/\\b\u0442\u0435\u043b\\.?:?/gi,' ').replace(/[\\s,;|]+/g,' ').trim();\n  contactNorm=(phones.join('; ')+(contactNorm?(' | '+contactNorm):'')).trim();\n}\nlet ch='unknown';const cl=contactNorm.toLowerCase();\nif(/t\\.me\\/|@[a-z0-9_]{4,}/.test(cl))ch='telegram';else if(/wa\\.me\\//.test(cl))ch='whatsapp';else if(/@[^\\s]+\\.[a-z]{2,}/.test(cl))ch='email';else if(phones.length)ch='phone';else if(/https?:\\/\\//.test(cl))ch='profile';else if(!contactNorm)ch='';\n\n// ---- guarantees / CTA heuristics ----\nfunction firstMatch(re){var mm=txt.match(re);return mm?cut(mm[0].replace(/\\s+/g,' '),120):'';}\nconst guarantees=firstMatch(/[^.\\n]*(\u0433\u0430\u0440\u0430\u043d\u0442\u0438|\u043f\u043e \u0434\u043e\u0433\u043e\u0432\u043e\u0440\u0443|\u0432\u0435\u0440\u043d\u0435\u043c|\u0432\u043e\u0437\u0432\u0440\u0430\u0442 \u0441\u0440\u0435\u0434\u0441\u0442\u0432|\u043e\u0444\u0438\u0446\u0438\u0430\u043b\u044c\u043d)[^.\\n]{0,80}/i);\nconst cta=firstMatch(/(\u043e\u0441\u0442\u0430\u0432\u044c\u0442\u0435 \u0437\u0430\u044f\u0432\u043a\u0443|\u043e\u0441\u0442\u0430\u0432\u0438\u0442\u044c \u0437\u0430\u044f\u0432\u043a\u0443|\u043f\u043e\u043b\u0443\u0447\u0438\u0442\u044c|\u043e\u0444\u043e\u0440\u043c\u0438\u0442\u044c|\u0437\u0430\u043a\u0430\u0437\u0430\u0442\u044c \u0437\u0432\u043e\u043d\u043e\u043a|\u043f\u043e\u0437\u0432\u043e\u043d\u0438\u0442\u0435|\u0443\u0437\u043d\u0430\u0442\u044c|\u0440\u0430\u0441\u0441\u0447\u0438\u0442\u0430\u0442\u044c)[^.\\n]{0,60}/i);\n\n// ---- offer field: clean text only; raw markdown preserved in audit ----\nconst offer_summary=offerIsRaw?cut(plainText(offerRaw),200):cut(offerRaw,300);\nconst offer_text_raw_audit=offerIsRaw?cut(offerRaw,2000):'';\n\n// ---- evidence-based confidence (S2-D11) ----\nlet conf=25;\nif(brand_source==='analysis')conf+=20;else if(brand_source==='title'||brand_source==='domain')conf+=8;\nif(offer_summary)conf+=12;\nif(str(r.terms))conf+=12;\nif(service_primary!=='unknown')conf+=12;\nif(contactNorm)conf+=8;\nif(txt.length>=400)conf+=10;\nif(page_type!=='other')conf+=5;\nif(fallbackParse||offerIsRaw)conf=Math.min(conf,40);\nif(broken_brand)conf=Math.min(conf,35);\nconf=Math.max(5,Math.min(95,conf));\n\n// ---- quality status / report eligibility / fallback reason ----\nlet quality_status='healthy';\nlet fallback_reason='';\nif(page_type==='other'&&!title&&!txt){quality_status='quarantined';fallback_reason='no page evidence (title/content empty)';}\nelse if(fallbackParse){quality_status='degraded';fallback_reason='primary+repair parse failed; deterministic fallback';}\nelse if(offerIsRaw){quality_status='degraded';fallback_reason='offer was raw markdown; preserved in audit only';}\nelse if(broken_brand){quality_status='degraded';fallback_reason='no brand evidence in title/domain/url';}\nconst report_eligible=(quality_status==='healthy');\n\n// ---- cost telemetry (S2-D17): unknown != zero ----\nconst primary_calls=1;\nconst repair_calls=(r.repair_used===true||low(r.repair_status)&&low(r.repair_status)!=='')?1:0;\nfunction numOrNull(v){var n=Number(v);return isFinite(n)&&str(v)!==''?n:null;}\nconst input_tokens=numOrNull(r.input_tokens);\nconst output_tokens=numOrNull(r.output_tokens);\nconst cost_status='unknown'; // Firecrawl/Claude $ not recovered into the row yet\n\n// --- Stage C Closure Patch 3: snapshot-level quality accounting (degraded/quarantined/snapshots_written) ---\n(function(){var __sd=$getWorkflowStaticData('global');var __r=(__sd.wf04_run=__sd.wf04_run||{});__r.snapshots_written=(__r.snapshots_written||0)+1;if(quality_status==='degraded')__r.degraded=(__r.degraded||0)+1;else if(quality_status==='quarantined')__r.quarantined=(__r.quarantined||0)+1;})();\nconst stamp=moscowStamp();\nreturn [{ json: {\n  snapshot_id:'site_snap_'+stamp+'_'+(Number(r.batch_index)||0),\n  created_at:moscowIsoNow(),\n  run_id:str(r.run_id)||('wf04_'+stamp),\n  source_run_id:str(r.run_id),\n  workflow_run_id:'wf04_'+stamp,\n  data_mode:'live',\n  agent_request_id:str(r.run_id),\n  source_url:src,\n  domain:domain,\n  company_name:cut(company,120),\n  brand_source:brand_source,\n  broken_brand:broken_brand,\n  page_type:page_type,\n  title:cut(title,160),\n  offer_summary:offer_summary,\n  offer_text_raw_audit:offer_text_raw_audit,\n  prices_terms:cut(str(r.terms),300),\n  guarantees:guarantees,\n  service_types:services.join(', '),\n  service_primary:service_primary,\n  service_secondary:service_secondary,\n  detected_pains:cut(str(r.detected_need),200),\n  contact_public:sheetSafe(contactNorm),\n  contact_public_raw_audit:sheetSafe(contactRaw),\n  contact_channel:ch,\n  cta_text:cta,\n  content_hash:hash(txt),\n  change_type:'baseline',\n  previous_snapshot_id:'',\n  source_confidence:conf,\n  quality_status:quality_status,\n  report_eligible:report_eligible,\n  fallback_reason:fallback_reason,\n  parse_method:str(r.parse_method),\n  repair_used:r.repair_used===true,\n  primary_calls:primary_calls,\n  repair_calls:repair_calls,\n  firecrawl_calls:1,\n  claude_calls:primary_calls+repair_calls,\n  input_tokens:input_tokens,\n  output_tokens:output_tokens,\n  actual_source_cost_usd:null,\n  actual_llm_cost_usd:null,\n  cost_status:cost_status,\n  notes:'wf04 snapshot (Stage C Closure Patch 2); brand_source='+brand_source+'; offer/prices from analysis, raw markdown kept in audit; confidence evidence-based; change_type diff is Phase C'\n}}];"
      }
    },
    {
      "id": "wf04-app-snap",
      "name": "Append competitor_site_snapshots",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        1720,
        1040
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "competitor_site_snapshots",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "wf04-build-runlog",
      "name": "Build live_source_runs Row",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        300,
        160
      ],
      "parameters": {
        "jsCode": "function moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nfunction moscowStamp(){var z=function(n){return String(n).padStart(2,'0');};var m=new Date(Date.now()+10800000);return m.getUTCFullYear()+z(m.getUTCMonth()+1)+z(m.getUTCDate())+'_'+z(m.getUTCHours())+z(m.getUTCMinutes())+z(m.getUTCSeconds());}\n// WF04 live_source_runs ledger. CANONICAL IDENTITY (Stage 3 closure, DEC-147): the connector's canonical\n// source_run_id (Set URL List run_id = 'firecrawl_<stamp>') is the join key used by raw records, snapshots,\n// source_health and reports. The workflow-local execution id ('wf04_<stamp>') is recorded SEPARATELY as\n// workflow_run_id and NEVER replaces source_run_id (the pre-fix wf04_* run_id broke WF16's join). One row\n// per run; aggregates the SplitInBatches DONE output via $input.all().\nfunction str(v){return v==null?'':String(v).trim();}function low(v){return str(v).toLowerCase();}\nconst items=$input.all().map(i=>i.json);\nconst total=items.length;\nconst written=items.filter(r=>str(r.note)==='processed_by_workflow_04').length; // non-duplicate registry rows\nconst dup=total-written; // duplicates routed to skipped_log (no Firecrawl call)\nlet sourceRunId='';try{sourceRunId=str($('Set URL List').first().json.run_id);}catch(e){}\nif(!sourceRunId)sourceRunId='firecrawl_'+moscowStamp();\nconst workflowRunId='wf04_'+sourceRunId.replace(/^firecrawl_/,'');\nlet agentRequestId='';try{agentRequestId=str($('Set URL List').first().json.agent_request_id);}catch(e){}\nlet received=total;try{received=$('Set URL List').all().length||total;}catch(e){}\n// run-level call accounting from staticData (primary/repair separated, S2-D7/D8/D16)\nconst __sd=$getWorkflowStaticData('global');const __r=(__sd&&__sd.wf04_run)||{};\nconst firecrawlCalls=Number(__r.firecrawl_calls)||written;\nconst primaryCalls=Number(__r.primary_calls)||written;\nconst repairCalls=Number(__r.repair_calls)||0;\nconst claudeCalls=Number(__r.claude_calls)||(primaryCalls+repairCalls);\nconst externalCalls=firecrawlCalls; // paid source (Firecrawl) calls\n// SOURCE-REUSE-001: reuse accounting + a stash for the (now downstream) agent_requests/summary builders,\n// which no longer receive the loop DONE items directly (the done-side is serialized for a deterministic return).\nconst reused=Number(__r.reused)||0;\nconst reuseFailed=Number(__r.reuse_failed)||0;\nconst pureReuse=(written===0&&reused>0&&firecrawlCalls===0);\n__sd.wf04_run=Object.assign({},__r,{done_total:total,registry_rows_written:written});\n// \u00a712 cost: a paid external/LLM call whose provider cost was NOT recovered is 'unknown' (never 0 / not_applicable)\nconst sourceCostStatus=(externalCalls>0?'unknown':'not_applicable');\nconst llmCostStatus=(claudeCalls>0?'unknown':'not_applicable');\nconst costStatus=((sourceCostStatus==='unknown'||llmCostStatus==='unknown')?'unknown':'not_applicable');\nreturn [{ json: {\n  run_id:sourceRunId,            // backward-compat alias == canonical source_run_id (join-safe)\n  source_run_id:sourceRunId,     // CANONICAL join key (raw/snapshot/source_health/report)\n  workflow_run_id:workflowRunId, // workflow-local execution id (was silently used as run_id)\n  agent_request_id:agentRequestId,\n  created_at:moscowIsoNow(),\n  workflow:'WF04 - Firecrawl URL List Resilient',\n  source_family:'web_competitor',\n  platform:'website',\n  mode:(pureReuse?'reuse':'live'),\n  data_mode:'live',\n  approval_token_used:'not_required', // WF04 has no paid-approval token gate (Firecrawl url list); accurate semantic (\u00a72.7)\n  source_allowlist:'operator url list (Set URL List, <=5)',\n  max_items:5,\n  items_received:received,\n  items_relevant:written,\n  items_written_raw:written,\n  items_unique:written,\n  items_duplicate:dup,\n  hard_skipped:0,\n  external_calls:externalCalls,\n  firecrawl_calls:firecrawlCalls,\n  primary_calls:primaryCalls,\n  repair_calls:repairCalls,\n  estimated_source_cost_usd:0,\n  actual_source_cost_usd:(externalCalls===0?0:null),\n  source_cost_status:sourceCostStatus,\n  llm_calls:claudeCalls,\n  estimated_llm_cost_usd:0,\n  actual_llm_cost_usd:(claudeCalls===0?0:null),\n  llm_cost_status:llmCostStatus,\n  cost_status:costStatus,\n  status:'completed',\n  error_summary:'',\n  operator_next_action:'WF16 (score source_health by source_run_id) -> WF08 (analyze canonical records) -> WF10/WF12 report.',\n  notes:(pureReuse?('REUSE run: no new collection; rows re-emitted from accepted snapshot of source_run '+String(__r.original_snapshot_run_id||'')+'; health inherited ($0). '):'')+'Web competitor scrape. Canonical source_run_id='+sourceRunId+' (= raw/snapshot/source_health join key); workflow_run_id='+workflowRunId+'. received=Set URL List; written=non-duplicate registry rows. Firecrawl/Claude $ not recovered => cost_status=unknown; primary/repair calls separated from staticData.'\n}}];"
      }
    },
    {
      "id": "wf04-app-runlog",
      "name": "Append live_source_runs",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        520,
        160
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "live_source_runs",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "wf04-build-ar",
      "name": "Build agent_requests Row",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        300,
        360
      ],
      "parameters": {
        "jsCode": "function moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nfunction moscowStamp(){var z=function(n){return String(n).padStart(2,'0');};var m=new Date(Date.now()+10800000);return m.getUTCFullYear()+z(m.getUTCMonth()+1)+z(m.getUTCDate())+'_'+z(m.getUTCHours())+z(m.getUTCMinutes())+z(m.getUTCSeconds());}\n// WF04 agent_requests row (21 cols). One per run, aggregated from the DONE output via $input.all().\nfunction str(v){return v==null?'':String(v).trim();}\n// SOURCE-REUSE-001: the done-side is serialized (this node now follows Append live_source_runs), so the loop\n// DONE counts come from the staticData stash written by Build live_source_runs Row.\nconst __sd=$getWorkflowStaticData('global');const __r=(__sd&&__sd.wf04_run)||{};\nconst total=Number(__r.done_total)||0;\nconst written=Number(__r.registry_rows_written)||0;\nconst reused=Number(__r.reused)||0;\nconst dup=Math.max(0,total-written-reused-(Number(__r.reuse_failed)||0));\nlet stamp='';try{stamp=str($('Set URL List').first().json.run_id).replace(/^firecrawl_/,'');}catch(e){stamp=String(Date.now());}\nreturn [{ json: {\n  agent_request_id:'wf04_'+stamp,\n  created_at:moscowIsoNow(),\n  requested_by:'operator',\n  request_text:'Firecrawl competitor website scrape (approved URL list, <=5).',\n  request_type:'web_competitor_scrape',\n  source_scope:'competitor_websites',\n  platforms:'website',\n  query:'',\n  region:'\u041c\u043e\u0441\u043a\u0432\u0430/\u041c\u041e',\n  service_focus:'credit_broker',\n  requested_limit:5,\n  status:'completed',\n  plan_summary:'WF04 Firecrawl URL list -> business route sheets + url_registry + competitor_site_snapshots.',\n  estimated_source_cost_usd:0,\n  estimated_analysis_cost_usd:0,\n  approval_required:false,\n  approved_by:'',\n  approved_at:'',\n  result_summary:'urls_processed='+total+'; scraped_written='+written+'; reused='+reused+'; reuse_failed='+(Number(__r.reuse_failed)||0)+'; duplicates='+dup+'; snapshots appended for non-error competitor pages.',\n  next_action:(written>0?'Review competitor_site_snapshots + monitor_queue; refresh WF12.':(reused>0?'Reused accepted snapshot(s); report built from saved data at $0 collection cost.':'No new pages written (all duplicates/errors); inspect Firecrawl output.')),\n  notes:'Firecrawl + Claude per non-duplicate URL. No outreach. Cost not recovered in-row (see Firecrawl/Anthropic dashboards).'\n}}];"
      }
    },
    {
      "id": "wf04-app-ar",
      "name": "Append agent_requests",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        520,
        360
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "agent_requests",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "parameters": {
        "jsCode": "// WF04 Final Summary Output (Stage C Closure Patch 2, S2-D7). Aggregates run-level accounting from\n// staticData (written by the parse + Normalize+Route nodes) + the SplitInBatches DONE output. $0, no calls.\nfunction str(v){return v==null?'':String(v).trim();}\nconst __sd=$getWorkflowStaticData('global');const a=__sd.wf04_run||{};\n// SOURCE-REUSE-001: serialized done-side \u2014 counts come from the staticData stash (Build live_source_runs Row).\nconst registryWritten=Number(a.registry_rows_written)||0;\nlet runId='';try{runId='wf04_'+str($('Set URL List').first().json.run_id).replace(/^firecrawl_/,'');}catch(e){runId='wf04_unknown';}\nlet urlsReceived=Number(a.urls_received)||0;try{urlsReceived=urlsReceived||$('Set URL List').all().length;}catch(e){}\nconst primaryCalls=Number(a.primary_calls)||0,repairCalls=Number(a.repair_calls)||0;\nconst primaryOk=Number(a.primary_parse_success)||0,primaryFail=Number(a.primary_parse_failure)||0;\nconst repairOk=Number(a.repair_success)||0,repairFail=Number(a.repair_failure)||0;\nconst detFallback=Number(a.deterministic_fallback)||0;\nconst claudeCalls=Number(a.claude_calls)||(primaryCalls+repairCalls);\nconst firecrawlCalls=Number(a.firecrawl_calls)||Number(a.urls_scraped)||0;\nconst repairRate=primaryCalls>0?Math.round((repairCalls/primaryCalls)*1000)/10:0;\nconst errors=[];if(primaryFail>repairOk+detFallback)errors.push('unrecovered_primary_parse_failures');\nconst __reused=Number(a.reused)||0;\nconst __reuseFailed=Number(a.reuse_failed)||0;\nconst __force=(a.force_reprocess===true);\nconst __pureReuse=(registryWritten===0&&__reused>0&&firecrawlCalls===0);\nconst __execMode=(__pureReuse?'reuse':(__reused>0?'mixed':(__force?'refresh':'collect')));\nconst __outcome=(__pureReuse?'reused_snapshot':(registryWritten>0?(__force?'refreshed_with_data':'collected_with_data'):(__reuseFailed>0?'empty_response':'no_relevant_content')));\nreturn [{ json: {\n  workflow:'04 - Firecrawl URL List Resilient',\n  run_id:runId,\n  agent_request_id:runId,\n  urls_received:urlsReceived,\n  urls_scraped:Number(a.urls_scraped)||0,\n  primary_calls:primaryCalls,\n  primary_parse_successes:primaryOk,\n  primary_parse_failures:primaryFail,\n  repair_calls:repairCalls,\n  repair_successes:repairOk,\n  repair_failures:repairFail,\n  deterministic_fallback_count:detFallback,\n  repair_rate_pct:repairRate,\n  degraded_count:Number(a.degraded)||0,\n  quarantined_count:Number(a.quarantined)||0,\n  snapshots_written:Number(a.snapshots_written)||registryWritten,\n  registry_rows_written:registryWritten,\n  monitor_rows_written:registryWritten,\n  firecrawl_calls:firecrawlCalls,\n  claude_calls:claudeCalls,\n  input_tokens:null,\n  output_tokens:null,\n  actual_source_cost_usd:(firecrawlCalls===0?0:null),\n  actual_llm_cost_usd:(claudeCalls===0?0:null),\n  cost_status:((firecrawlCalls===0&&claudeCalls===0)?'not_applicable':'unknown'),\n  technical_errors:errors.join('; '),\n  status:'completed',\n  // SOURCE-REUSE-001: the adapter contract WF20 consumes \u2014 typed execution mode + outcome + reuse lineage.\n  items_received:urlsReceived,\n  items_written:registryWritten+__reused,\n  items_unique:registryWritten+__reused,\n  external_calls:firecrawlCalls,\n  execution_mode:__execMode,\n  source_outcome:__outcome,\n  reused_count:__reused,\n  reuse_failed_count:__reuseFailed,\n  original_snapshot_run_id:str(a.original_snapshot_run_id),\n  original_snapshot_collected_at:str(a.original_snapshot_collected_at),\n  next_action:(repairRate>=40?('HIGH repair rate ('+repairRate+'%) \u2014 treat affected snapshots as degraded; inspect Firecrawl/Claude output quality before trusting the report.'):(__execMode==='reuse'?('Report built from the accepted snapshot of source_run '+str(a.original_snapshot_run_id)+' (collected '+str(a.original_snapshot_collected_at)+'); $0 collection.'):'Review competitor_site_snapshots + monitor_queue; refresh WF12 report.'))\n}}];"
      },
      "id": "aa040000-0000-0000-0000-0000000000fs",
      "name": "Final Summary Output",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        240,
        980
      ]
    },
    {
      "id": "wf04-build-canonical-raw",
      "name": "Build Canonical Raw Record",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1980,
        820
      ],
      "parameters": {
        "jsCode": "// --- PARSE-OUTCOME-001: embedded n8n/lib/parse_outcome.js (drift-proof; test asserts equality) ---\n// parse_outcome.js \u2014 PARSE-OUTCOME-001. The ONE mapping from \"how did the LLM parse go?\" to a quality verdict.\n//\n// OPERATOR DECISION (2026-07-17): a single successful bounded repair is an ACCEPTABLE result when \u2014 and only\n// when \u2014 the repaired payload passed full local schema validation, required fields are present, evidence\n// references are valid, structural/semantic rules pass, no forbidden data is present, and no second repair was\n// needed. Such a result is marked repaired / accepted_with_repair, audited, and CONFIDENCE-CAPPED \u2014 but it is NOT\n// placed into pending human review and NOT quarantined merely because a repair was used.\n//\n// Why this existed as a defect: WF04 stamped `repair_used === true` -> quality_status='degraded' ->\n// review_status='pending'. Live (autolombardn1.ru, WF04 exec 929) the repair SUCCEEDED and produced excellent data\n// (company, full offer text, service, valid evidence url) \u2014 yet the record was degraded+pending, so WF16 derived\n// report_candidate=false, raised the CRITICAL `no_detail_records` flag, quarantined a score-81 run, and WF10\n// dropped it. The report was empty and WF28 never ran. Meanwhile WF28 (Stage F) already treats a successful repair\n// as shippable (`quality_status='repaired'`, enriched=true, live-proven exec 834). Two of our own contracts\n// disagreed; this lib is the single place they now agree.\n//\n// A FAILED repair, invalid evidence, semantic failure, malformed result or fallback-only result stays FAIL-CLOSED.\n//\n// Embeddable: unique po*-prefixed names, no cross-lib require.\n\nfunction poStr(v) { return v == null ? '' : String(v); }\nfunction poLow(v) { return poStr(v).trim().toLowerCase(); }\nfunction poBool(v) { return v === true || poLow(v) === 'true'; }\n\n// The canonical parse outcomes. Anything else is unknown -> treated as invalid (fail-closed).\nvar PARSE_OUTCOMES = {\n  PRIMARY_VALID: 'primary_valid',                 // first parse validated cleanly\n  REPAIRED_VALID: 'repaired_valid',               // ONE repair, payload fully validated\n  DETERMINISTIC_FALLBACK: 'deterministic_fallback', // no successful LLM parse; deterministic facts only\n  INVALID: 'invalid',                             // malformed / failed validation / repair failed\n  PROVIDER_FAILED: 'provider_failed'              // transport/provider error; nothing was parsed\n};\n// Quality verdicts this lib may assign. 'accepted_with_repair' is the new, audited middle state.\nvar PO_QUALITY = {\n  HEALTHY: 'healthy',\n  ACCEPTED_WITH_REPAIR: 'accepted_with_repair',\n  DEGRADED: 'degraded',\n  QUARANTINED: 'quarantined'\n};\n// A repaired result is trustworthy enough to report, but never as trustworthy as a clean primary parse. ONE\n// documented rule: cap it. (The evidence is identical; only our confidence in the extraction is lower.)\nvar PO_REPAIRED_CONFIDENCE_CAP = 75;\n// The contract is ONE bounded repair. A second repair is never \"accepted_with_repair\".\nvar PO_MAX_REPAIRS = 1;\n\n// Quality verdicts that may enter a report / feed Claude.\nfunction poIsAccepted(q) { return q === PO_QUALITY.HEALTHY || q === PO_QUALITY.ACCEPTED_WITH_REPAIR; }\nfunction poIsRepaired(q) { return q === PO_QUALITY.ACCEPTED_WITH_REPAIR; }\n\n// classifyParseOutcome(rec) -> one of PARSE_OUTCOMES.\n// rec: { processing_status, parse_method, repair_used, repair_status, repair_count, validation_ok,\n//        evidence_valid, has_evidence }\nfunction classifyParseOutcome(rec) {\n  rec = rec || {};\n  var ps = poLow(rec.processing_status);\n  var pm = poLow(rec.parse_method);\n  var repairStatus = poLow(rec.repair_status);\n  var repairUsed = poBool(rec.repair_used);\n  var repairs = Number(rec.repair_count);\n  if (!isFinite(repairs)) repairs = repairUsed ? 1 : 0;\n\n  if (ps === 'provider_failed' || pm === 'firecrawl_error' || poLow(rec.access_failure) === 'true') return PARSE_OUTCOMES.PROVIDER_FAILED;\n  if (ps === 'technical_error') return PARSE_OUTCOMES.INVALID;\n  // an explicit deterministic fallback never claims a successful LLM parse\n  if (pm === 'deterministic_competitor_fallback' || pm === 'deterministic_fallback' || repairStatus === 'failed_fallback') {\n    return PARSE_OUTCOMES.DETERMINISTIC_FALLBACK;\n  }\n  // hard local gates \u2014 these fail closed regardless of how the parse got here\n  if (rec.validation_ok === false || rec.evidence_valid === false) return PARSE_OUTCOMES.INVALID;\n  if (rec.has_evidence === false) return PARSE_OUTCOMES.INVALID;\n  if (repairUsed || pm === 'repaired_json') {\n    if (repairStatus === 'failed' || repairs > PO_MAX_REPAIRS) return PARSE_OUTCOMES.INVALID; // >1 repair is never accepted\n    return PARSE_OUTCOMES.REPAIRED_VALID;\n  }\n  if (ps === 'parsed_success' || pm === 'primary_json' || pm === '') return PARSE_OUTCOMES.PRIMARY_VALID;\n  return PARSE_OUTCOMES.PRIMARY_VALID;\n}\n\n// qualityForOutcome(outcome, opts) -> { quality_status, review_status, report_candidate, flags[],\n//                                       confidence_cap, repair_used, repair_success }\n// `report_candidate` here is the PARSE dimension only \u2014 business relevance, access status and evidence quality are\n// decided elsewhere and still apply. Repair status alone must never decide relevance or source health.\nfunction qualityForOutcome(outcome, opts) {\n  opts = opts || {};\n  var isCompetitor = opts.is_competitor !== false;\n  switch (outcome) {\n    case PARSE_OUTCOMES.PRIMARY_VALID:\n      return { quality_status: PO_QUALITY.HEALTHY, review_status: 'confirmed', report_candidate: isCompetitor,\n        flags: [], confidence_cap: null, repair_used: false, repair_success: false };\n    case PARSE_OUTCOMES.REPAIRED_VALID:\n      // audited + capped, but reportable: the payload passed the SAME local validation as a primary parse.\n      return { quality_status: PO_QUALITY.ACCEPTED_WITH_REPAIR, review_status: 'confirmed', report_candidate: isCompetitor,\n        flags: ['repaired_parse'], confidence_cap: PO_REPAIRED_CONFIDENCE_CAP, repair_used: true, repair_success: true };\n    case PARSE_OUTCOMES.DETERMINISTIC_FALLBACK:\n      // deterministic facts may still ship, but we never claim a successful enrichment.\n      return { quality_status: PO_QUALITY.DEGRADED, review_status: 'pending', report_candidate: false,\n        flags: ['deterministic_fallback'], confidence_cap: 50, repair_used: false, repair_success: false };\n    case PARSE_OUTCOMES.PROVIDER_FAILED:\n      return { quality_status: PO_QUALITY.QUARANTINED, review_status: 'pending', report_candidate: false,\n        flags: ['provider_failed'], confidence_cap: 10, repair_used: false, repair_success: false };\n    default: // INVALID\n      return { quality_status: PO_QUALITY.QUARANTINED, review_status: 'pending', report_candidate: false,\n        flags: ['parse_invalid'], confidence_cap: 10, repair_used: false, repair_success: false };\n  }\n}\n\n// Apply the documented confidence rule. A repaired result is capped, never boosted.\nfunction poCapConfidence(score, outcome) {\n  var n = Number(score); if (!isFinite(n)) n = 0;\n  var q = qualityForOutcome(outcome, {});\n  if (q.confidence_cap == null) return n;\n  return Math.min(n, q.confidence_cap);\n}\n// --- end embedded parse_outcome ---\n\n// WF04 -> canonical raw_market_records (Stage 3 closure, DEC-148). WF04 is the WEBSITE SOURCE ADAPTER:\n// transport (Firecrawl) + cleaning + source metadata/evidence. It emits ONE canonical, WF16/WF08-compatible\n// raw record per successfully scraped URL with analysis_status='pending'. WF08 is the SINGLE semantic owner\n// and consumes this record exactly once; WF04's own extraction is kept here as SOURCE HINTS/evidence\n// (service_hint/competitor_name/offer_text), never as final analyzer authority. No external calls.\nfunction str(v){return v==null?'':String(v).trim();}\nfunction low(v){return str(v).toLowerCase();}\nfunction cut(s,n){s=str(s);return s.length>n?s.slice(0,n):s;}\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nfunction moscowStamp(){var z=function(n){return String(n).padStart(2,'0');};var m=new Date(Date.now()+10800000);return m.getUTCFullYear()+z(m.getUTCMonth()+1)+z(m.getUTCDate())+'_'+z(m.getUTCHours())+z(m.getUTCMinutes())+z(m.getUTCSeconds());}\nfunction domainOf(u){u=str(u).toLowerCase();var m=u.match(/^https?:\\/\\/([^\\/?#]+)/);var h=m?m[1]:'';return h.replace(/^www\\./,'');}\nfunction brandFromDomain(d){d=str(d).replace(/^www\\./,'').split('.')[0];if(!d)return '';return d.charAt(0).toUpperCase()+d.slice(1);}\nconst r=$json;\nconst route=low(r.route);const ps=low(r.processing_status);const pm=low(r.parse_method);\nconst src=str(r.source_url);\nif(!src) return []; // no URL => nothing was scraped; no canonical record\nlet sourceRunId='';try{sourceRunId=str($('Set URL List').first().json.run_id);}catch(e){}\nif(!sourceRunId)sourceRunId=str(r.run_id);\nlet agentReq='';try{agentReq=str($('Set URL List').first().json.agent_request_id);}catch(e){}\nif(!agentReq)agentReq=str(r.agent_request_id)||('wf04_req_'+sourceRunId.replace(/^firecrawl_/,''));\nconst stamp=sourceRunId.replace(/^firecrawl_/,'')||moscowStamp();\nconst workflowRunId='wf04_'+stamp;\nconst idx=Number(r.batch_index)||0;\nconst sourceRecordId='wf04_rec_'+stamp+'_'+idx;\nconst domain=domainOf(src);\nconst text=cut(str(r.text_context),4000);\nlet company=str(r.company_name);\nif(!company||/^\u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442 \u0431\u0435\u0437 \u0431\u0440\u0435\u043d\u0434\u0430$/i.test(company)){var bd=brandFromDomain(domain);if(bd)company=bd;}\nconst title=company||domain;\n// quality gate (mirrors the snapshot logic): primary success=healthy; repaired/deterministic-fallback=degraded;\n// transport/parse failure or no page evidence=quarantined.\nconst hasEvidence=(text.length>=30)||str(r.offer_text)!=='';\nconst isCompetitor=(low(r.entity_type)==='competitor');\n// PARSE-OUTCOME-001 (operator decision 2026-07-17): the quality verdict comes from the CANONICAL contract, not\n// from a local conditional. A SINGLE successful bounded repair whose payload passed full local validation is\n// `accepted_with_repair` \u2014 audited + confidence-capped, but reportable. It is NOT degraded and NOT pending review\n// merely because a repair happened. This aligns WF04 with the already-proven WF28 contract, which has always\n// shipped a repaired analysis (live exec 834). A FAILED repair, a deterministic fallback, an invalid payload or a\n// provider failure all remain fail-closed exactly as before.\nvar __po=classifyParseOutcome({processing_status:ps,parse_method:pm,repair_used:r.repair_used,\n  repair_status:r.repair_status,repair_count:r.repair_calls,has_evidence:hasEvidence,\n  access_failure:(route==='technical_errors')?'true':''});\nvar __q=qualityForOutcome(__po,{is_competitor:isCompetitor});\nlet quality_status=__q.quality_status;var qflags=__q.flags.slice();\nif(ps==='technical_error'||route==='technical_errors'||!hasEvidence){quality_status='quarantined';qflags=['no_page_evidence'];}\nconst parse_outcome=__po;   // audited on the row; also surfaces via parse_method + quality_flags\n// report_candidate/eligibility follow the parse verdict, but relevance is decided independently upstream.\nconst report_eligible=(__q.report_candidate===true&&quality_status!=='quarantined'&&isCompetitor);\nconst review_status=(quality_status==='quarantined'?'pending':__q.review_status);\nconst recordTypeHint=isCompetitor?'competitor_activity':(low(r.entity_type)==='irrelevant'?'irrelevant':'market_signal');\nconst repair_calls=(r.repair_used===true||(low(r.repair_status)&&low(r.repair_status)!==''))?1:0;\n// Relevance signal (RELEV-WEB-001): populate canonical confidence_score (0-100) + content-derived\n// semantic_keywords + relevance_reason for every website record. Evidence is POST-LEVEL (extracted page\n// text/offer/terms only), never the URL alone. Mirrors WF11 confidence tiers (competitor>market_signal>irrelevant).\nvar SERVICE_TERMS=['\u043a\u0440\u0435\u0434\u0438\u0442\u043d\u044b\u0439 \u0431\u0440\u043e\u043a\u0435\u0440','\u0438\u043f\u043e\u0442\u0435\u0447\u043d\u044b\u0439 \u0431\u0440\u043e\u043a\u0435\u0440','\u0438\u043f\u043e\u0442\u0435\u043a','\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441','\u043a\u0440\u0435\u0434\u0438\u0442 \u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433','\u0437\u0430\u043b\u043e\u0433 \u043f\u0442\u0441','\u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e','\u0437\u0430\u043b\u043e\u0433 \u043d\u0435\u0434\u0432\u0438\u0436','\u043f\u043e\u0442\u0440\u0435\u0431\u0438\u0442\u0435\u043b\u044c\u0441\u043a','\u043a\u0440\u0435\u0434\u0438\u0442 \u043d\u0430\u043b\u0438\u0447\u043d\u044b\u043c\u0438','\u043f\u043e\u0441\u043b\u0435 \u043e\u0442\u043a\u0430\u0437','\u043f\u043b\u043e\u0445\u0430\u044f \u043a\u0440\u0435\u0434\u0438\u0442\u043d','\u043f\u043b\u043e\u0445\u043e\u0439 \u043a\u0440\u0435\u0434\u0438\u0442\u043d','\u043a\u0440\u0435\u0434\u0438\u0442\u043d\u0430\u044f \u0438\u0441\u0442\u043e\u0440\u0438','\u0434\u043b\u044f \u0431\u0438\u0437\u043d\u0435\u0441\u0430','\u043e\u0431\u043e\u0440\u043e\u0442\u043d','\u0442\u0435\u043d\u0434\u0435\u0440\u043d','\u0431\u0430\u043d\u043a\u043e\u0432\u0441\u043a \u0433\u0430\u0440\u0430\u043d\u0442','\u0430\u0432\u0442\u043e\u043a\u0440\u0435\u0434\u0438\u0442','\u0440\u0430\u0441\u0441\u0440\u043e\u0447\u043a','\u0437\u0430\u0439\u043c','\u043f\u043e\u0434\u0431\u043e\u0440 \u0431\u0430\u043d\u043a','\u043f\u043e\u0434\u0431\u043e\u0440 \u043a\u0440\u0435\u0434\u0438\u0442','\u0441\u043d\u0438\u0436\u0435\u043d \u0441\u0442\u0430\u0432\u043a','\u043e\u0434\u043e\u0431\u0440\u0435\u043d \u043a\u0440\u0435\u0434\u0438\u0442'];\nvar evBlob=low(text)+' '+low(str(r.offer_text))+' '+low(str(r.terms))+' '+low(str(r.service_type))+' '+low(str(r.detected_need));\nvar svcHits=[];for(var _si=0;_si<SERVICE_TERMS.length;_si++){var _t=SERVICE_TERMS[_si];if(evBlob.indexOf(_t)>=0&&svcHits.indexOf(_t)<0)svcHits.push(_t);}\nvar confidence_score;var relevance_reason;\nif(quality_status==='quarantined'||recordTypeHint==='irrelevant'){confidence_score=10;relevance_reason='\u043d\u0435\u0440\u0435\u043b\u0435\u0432\u0430\u043d\u0442\u043d\u043e: \u043d\u0435\u0442 \u0434\u043e\u043a\u0430\u0437\u0430\u0442\u0435\u043b\u044c\u043d\u043e\u0439 \u0431\u0430\u0437\u044b \u0441\u0442\u0440\u0430\u043d\u0438\u0446\u044b';}\nelse if(isCompetitor){var _base=(quality_status==='healthy'||quality_status==='accepted_with_repair')?70:50;confidence_score=poCapConfidence(Math.min(90,_base+svcHits.length*5),__po);relevance_reason='\u0441\u0430\u0439\u0442-\u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442 ('+(domain||'\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a')+'): '+(svcHits.slice(0,6).join(', ')||(low(str(r.service_type))||'\u043a\u0440\u0435\u0434\u0438\u0442\u043d\u044b\u0435 \u0443\u0441\u043b\u0443\u0433\u0438'));}\nelse{confidence_score=45;relevance_reason='\u0444\u0438\u043d\u0430\u043d\u0441\u043e\u0432\u044b\u0439 \u043a\u043e\u043d\u0442\u0435\u043a\u0441\u0442 (\u043d\u0435 \u043f\u0440\u044f\u043c\u043e\u0435 \u043f\u0440\u0435\u0434\u043b\u043e\u0436\u0435\u043d\u0438\u0435 \u0431\u0440\u043e\u043a\u0435\u0440\u0430)';}\nreturn [{ json: {\n  parse_outcome:parse_outcome, repair_used:__q.repair_used, repair_success:__q.repair_success,\n  record_id:sourceRecordId,\n  source_record_id:sourceRecordId,\n  agent_request_id:agentReq,\n  source_run_id:sourceRunId,\n  run_id:sourceRunId,\n  workflow_run_id:workflowRunId,\n  data_mode:'live',\n  created_at:str(r.created_at)||moscowIsoNow(),\n  source_type:'scraped_web',\n  platform:'website',\n  source_url:src,\n  post_url:src,\n  exact_evidence_url:(quality_status!=='quarantined'),\n  profile_url:'',\n  profile_name:company,\n  author_handle:'',\n  published_at:str(r.published_at),\n  region_hint:str(r.region),\n  service_hint:low(r.service_type)||'unknown',\n  title:title,\n  description:cut(text,800),\n  text_context:text,\n  comment_text:'',\n  contact_public:str(r.contact_public),\n  dedup_key:low(r.platform||'website')+'::scraped_web::'+(domain||sourceRecordId),\n  record_type_hint:recordTypeHint,\n  touchpoint_type:'competitor_website',\n  is_detail:true,\n  detail_fetch_required:false,\n  placeholder_title:false,\n  is_valid_listing:(quality_status!=='quarantined'),\n  competitor_related:isCompetitor,\n  competitor_name:isCompetitor?company:'',\n  interest_topic:str(r.detected_need),\n  probable_need:str(r.detected_need),\n  offer_text:cut(str(r.offer_text),600),\n  terms:cut(str(r.terms),300),\n  semantic_keywords:svcHits.slice(0,8).join(', '),\n  confidence_score:confidence_score,\n  relevance_reason:relevance_reason,\n  ad_channel_hint:'website',\n  lead_intent_hint:'low',\n  urgency_hint:'low',\n  lead_temperature:'cold',\n  quality_status:quality_status,\n  report_eligible:report_eligible,\n  llm_eligible:report_eligible,\n  review_status:review_status,\n  analysis_status:'pending',\n  quality_flags:qflags.join('; '),\n  dedup_status:'unique',\n  approval_status:'new',\n  estimated_analysis_cost_usd:0,\n  firecrawl_calls:1,\n  primary_calls:1,\n  repair_calls:repair_calls,\n  actual_source_cost_usd:null,\n  cost_status:'unknown',\n  parse_method:str(r.parse_method),\n  manager_note:'WF04 website source record; analysis_status=pending; WF08 is the semantic owner.',\n  notes:'wf04_canonical_raw_record_stage3_closure',\n  taxonomy_version:'semantic-v2.0'\n}}];"
      }
    },
    {
      "id": "wf04-append-raw",
      "name": "Append raw_market_records",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        2200,
        820
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "raw_market_records",
          "mode": "name"
        },
        "options": {},
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        }
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "parameters": {
        "inputSource": "workflowInputs",
        "workflowInputs": {
          "values": [
            {
              "name": "agent_request_id",
              "type": "string"
            },
            {
              "name": "source_run_id",
              "type": "string"
            },
            {
              "name": "workflow_run_id",
              "type": "string"
            },
            {
              "name": "data_mode",
              "type": "string"
            },
            {
              "name": "urls",
              "type": "array"
            },
            {
              "name": "force_reprocess",
              "type": "string"
            },
            {
              "name": "owner_user_id",
              "type": "string"
            },
            {
              "name": "source_execution_mode",
              "type": "string"
            },
            {
              "name": "freshness_days",
              "type": "string"
            }
          ]
        }
      },
      "type": "n8n-nodes-base.executeWorkflowTrigger",
      "typeVersion": 1.1,
      "position": [
        120,
        500
      ],
      "id": "wf04-subtrig",
      "name": "When Called by Agent"
    },
    {
      "id": "fc000000-0000-0000-0000-0000000000r1",
      "name": "Source Health Lookup",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        990,
        500
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "read",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "source_health",
          "mode": "name"
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "fc000000-0000-0000-0000-0000000000r2",
      "name": "Read Reuse Route Rows",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        1100,
        80
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "read",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "={{ $json.reuse_route }}",
          "mode": "name"
        },
        "filtersUI": {
          "values": [
            {
              "lookupColumn": "source_url",
              "lookupValue": "={{ $('Evaluate Dedup').first().json.source_url }}"
            }
          ]
        },
        "options": {
          "returnAllMatches": "returnAllMatches"
        }
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "fc000000-0000-0000-0000-0000000000r3",
      "name": "Build Reuse Records",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1320,
        80
      ],
      "parameters": {
        "jsCode": "// SOURCE-REUSE-001: re-emit the ORIGINAL accepted route rows under the CURRENT request's lineage. The alias row\n// keeps the original content and parsed_at (when it was really collected) and is explicitly marked as a reuse \u2014\n// parse_method=reused_snapshot / processing_status=reused / quality_flags+=reused_snapshot \u2014 so nothing downstream\n// can mistake it for a fresh collection. No Firecrawl, no Claude, no new url_registry/raw/snapshot rows: $0.\nconst ctx = $('Evaluate Dedup').first().json;\nfunction str(v){return v==null?'':String(v).trim();}\nfunction low(v){return str(v).toLowerCase();}\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nconst now = moscowIsoNow();\nfunction normUrl(u){ return low(u).replace(/^https?:\\/\\//,'').replace(/^www\\./,'').replace(/\\/+$/,''); }\nconst key = normUrl(ctx.source_url);\nlet rows = [];\ntry { rows = $('Read Reuse Route Rows').all().map(i => i.json); } catch(e) { rows = []; }\nconst BAD_PS = ['technical_error', 'failed', 'error'];\nconst usable = rows.filter(function (r) {\n  if (!r) return false;\n  if (normUrl(r.source_url) !== key) return false;\n  if (str(r.run_id) !== str(ctx.original_run_id)) return false;      // exactly the run the registry promised\n  if (BAD_PS.indexOf(low(r.processing_status)) >= 0) return false;\n  if (low(r.quality_status) === 'quarantined') return false;\n  if (r.report_eligible === false || low(r.report_eligible) === 'false') return false;\n  return true;\n}).slice(0, 10);\nconst __sd = $getWorkflowStaticData('global'); __sd.wf04_run = __sd.wf04_run || {};\nif (!usable.length) {\n  // Registry promised an accepted snapshot but the queue rows are gone/unusable. Fail closed and say so \u2014 never\n  // fabricate data, never pretend the source was collected.\n  __sd.wf04_run.reuse_failed = (Number(__sd.wf04_run.reuse_failed) || 0) + 1;\n  return [{ json: {\n    created_at: now, source_type: ctx.source_type || 'scraped_web', platform: ctx.platform || 'website',\n    source_url: str(ctx.source_url), parsed_at: ctx.parsed_at || now, published_at: '', freshness_status: 'unknown',\n    entity_type: 'irrelevant', company_name: '', profile_name: '', profile_url: '', region: '',\n    service_type: 'unknown', offer_text: '', terms: '', contact_public: '', text_context: '', detected_need: '',\n    competitor_strength: 1, lead_signal_score: 1, content_idea_score: 1, quality_score: 1,\n    reason: 'url_registry points to accepted snapshot run ' + str(ctx.original_run_id) + ' but no usable rows were found in ' + str(ctx.reuse_route) + '; ask for a refresh to recollect.',\n    recommended_action: 'ignore', status: 'failed', processing_status: 'reuse_readback_empty',\n    parse_method: 'reused_snapshot', parse_error: '',\n    raw_response_preview: 'Reuse readback found no usable rows for run ' + str(ctx.original_run_id) + '.',\n    route: 'skipped_log', needs_manual_review: false, repair_used: false, repair_status: '',\n    run_id: str(ctx.run_id), batch_index: ctx.batch_index, sx_reuse_failed: true\n  }}];\n}\n__sd.wf04_run.reused = (Number(__sd.wf04_run.reused) || 0) + usable.length;\n__sd.wf04_run.original_snapshot_run_id = str(ctx.original_run_id);\n__sd.wf04_run.original_snapshot_collected_at = str(usable[0].parsed_at || ctx.original_collected_at);\nreturn usable.map(function (r) {\n  const c = Object.assign({}, r);\n  delete c.row_number; // sheets read metadata, not a column\n  const flags = str(r.quality_flags);\n  return { json: Object.assign(c, {\n    created_at: now,                                  // the reuse EVENT time; parsed_at stays the ORIGINAL collection time\n    run_id: str(ctx.run_id),                          // current-request lineage (WF10 isolation)\n    source_run_id: str(ctx.run_id),\n    batch_index: ctx.batch_index,\n    data_mode: 'live',\n    processing_status: 'reused',\n    parse_method: 'reused_snapshot',\n    parse_error: '',\n    raw_response_preview: 'Reused accepted snapshot (source_run ' + str(ctx.original_run_id) + ', collected ' + str(r.parsed_at || ctx.original_collected_at) + '); no new collection ($0).',\n    quality_flags: flags ? (flags + '; reused_snapshot') : 'reused_snapshot',\n    sx_reuse_failed: false\n  })};\n});"
      }
    },
    {
      "id": "fc000000-0000-0000-0000-0000000000r4",
      "name": "IF Reuse Records OK?",
      "type": "n8n-nodes-base.if",
      "typeVersion": 2,
      "position": [
        1540,
        80
      ],
      "parameters": {
        "conditions": {
          "options": {
            "caseSensitive": true,
            "leftValue": "",
            "typeValidation": "loose"
          },
          "conditions": [
            {
              "id": "sx-reuse-ok-01",
              "leftValue": "={{ $json.sx_reuse_failed === true }}",
              "rightValue": true,
              "operator": {
                "type": "boolean",
                "operation": "true",
                "singleValue": true
              }
            }
          ],
          "combinator": "and"
        }
      }
    },
    {
      "id": "fc000000-0000-0000-0000-0000000000r5",
      "name": "Append Reuse Route Row",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        1760,
        0
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "={{ $json.route }}",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "onError": "continueRegularOutput",
      "alwaysOutputData": true,
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "fc000000-0000-0000-0000-0000000000r6",
      "name": "Build Reuse Health Row",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1980,
        0
      ],
      "parameters": {
        "jsCode": "// SOURCE-REUSE-001: WF10 runs with require_source_health=true (production fail-closed), so the CURRENT run needs an\n// eligible source_health row or the reused rows are excluded all over again. The honest verdict for reused data is\n// the ORIGINAL run's verdict (WF16 scored that content when it was collected): inherit it under the current run's\n// lineage, marked reused_snapshot, with zero calls and zero cost. WF16 itself SKIPS pure-reuse runs (it would\n// otherwise score \"0 fresh records\" and overwrite this row in the last-wins eligibility index).\nconst ctx = $('Evaluate Dedup').first().json;\nfunction str(v){return v==null?'':String(v).trim();}\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nfunction moscowStamp(){var z=function(n){return String(n).padStart(2,'0');};var m=new Date(Date.now()+10800000);return m.getUTCFullYear()+z(m.getUTCMonth()+1)+z(m.getUTCDate())+'_'+z(m.getUTCHours())+z(m.getUTCMinutes())+z(m.getUTCSeconds());}\nconst h = ctx.original_health_row || {};\nconst c = Object.assign({}, h);\ndelete c.row_number;\nlet agentRequestId = '';\ntry { agentRequestId = str($('Set URL List').first().json.agent_request_id); } catch(e) {}\nreturn [{ json: Object.assign(c, {\n  quality_evaluation_id: 'qe_reuse_' + moscowStamp(),\n  evaluated_at: moscowIsoNow(),\n  source_run_id: str(ctx.run_id),\n  agent_request_id: agentRequestId,\n  external_calls: 0,\n  llm_calls: 0,\n  actual_source_cost_usd: 0,\n  actual_llm_cost_usd: 0,\n  source_cost_status: 'not_applicable',\n  llm_cost_status: 'not_applicable',\n  cost_per_unique_record: 0,\n  quality_flags: (str(h.quality_flags) ? str(h.quality_flags) + '; ' : '') + 'reused_snapshot',\n  notes: 'Reused accepted snapshot of source_run ' + str(ctx.original_run_id) + ' (collected ' + str(ctx.original_collected_at) + '); health inherited from the original run; no new collection ($0).'\n})}];"
      }
    },
    {
      "id": "fc000000-0000-0000-0000-0000000000r7",
      "name": "Append Reuse source_health",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        2200,
        0
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "={{ $env.MS_SPREADSHEET_ID || \"PASTE_SPREADSHEET_ID\" }}",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "source_health",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "onError": "continueRegularOutput",
      "alwaysOutputData": true,
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    }
  ],
  "connections": {
    "Manual Start": {
      "main": [
        [
          {
            "node": "Set URL List",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Set URL List": {
      "main": [
        [
          {
            "node": "Loop Over Items",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Loop Over Items": {
      "main": [
        [
          {
            "node": "Build live_source_runs Row",
            "type": "main",
            "index": 0
          }
        ],
        [
          {
            "node": "Normalize URL for Dedup",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Normalize URL for Dedup": {
      "main": [
        [
          {
            "node": "Registry Lookup",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Registry Lookup": {
      "main": [
        [
          {
            "node": "Source Health Lookup",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Evaluate Dedup": {
      "main": [
        [
          {
            "node": "IF Reuse?",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Append Skipped Log (Duplicate)": {
      "main": [
        [
          {
            "node": "Loop Over Items",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build Firecrawl Request": {
      "main": [
        [
          {
            "node": "Firecrawl Scrape API",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Firecrawl Scrape API": {
      "main": [
        [
          {
            "node": "Normalize Firecrawl Output",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Normalize Firecrawl Output": {
      "main": [
        [
          {
            "node": "IF Firecrawl Normalized OK?",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "IF Firecrawl Normalized OK?": {
      "main": [
        [
          {
            "node": "Build Primary Claude Request",
            "type": "main",
            "index": 0
          }
        ],
        [
          {
            "node": "Append to Dynamic Route Sheet",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build Primary Claude Request": {
      "main": [
        [
          {
            "node": "Claude Primary API Request",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Claude Primary API Request": {
      "main": [
        [
          {
            "node": "Parse Primary JSON",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Parse Primary JSON": {
      "main": [
        [
          {
            "node": "IF Primary Parse OK?",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "IF Primary Parse OK?": {
      "main": [
        [
          {
            "node": "Normalize + Route",
            "type": "main",
            "index": 0
          }
        ],
        [
          {
            "node": "Build Repair Request",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build Repair Request": {
      "main": [
        [
          {
            "node": "Claude Repair API Request",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Claude Repair API Request": {
      "main": [
        [
          {
            "node": "Parse Repaired JSON",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Parse Repaired JSON": {
      "main": [
        [
          {
            "node": "Normalize + Route",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Normalize + Route": {
      "main": [
        [
          {
            "node": "Append to Dynamic Route Sheet",
            "type": "main",
            "index": 0
          },
          {
            "node": "Build competitor_site_snapshots Row",
            "type": "main",
            "index": 0
          },
          {
            "node": "Build Canonical Raw Record",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Append to Dynamic Route Sheet": {
      "main": [
        [
          {
            "node": "Build Registry Row",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build Registry Row": {
      "main": [
        [
          {
            "node": "Append url_registry",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Append url_registry": {
      "main": [
        [
          {
            "node": "Loop Over Items",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build competitor_site_snapshots Row": {
      "main": [
        [
          {
            "node": "Append competitor_site_snapshots",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build live_source_runs Row": {
      "main": [
        [
          {
            "node": "Append live_source_runs",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build agent_requests Row": {
      "main": [
        [
          {
            "node": "Append agent_requests",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build Canonical Raw Record": {
      "main": [
        [
          {
            "node": "Append raw_market_records",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "When Called by Agent": {
      "main": [
        [
          {
            "node": "Set URL List",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "IF Reuse?": {
      "main": [
        [
          {
            "node": "Read Reuse Route Rows",
            "type": "main",
            "index": 0
          }
        ],
        [
          {
            "node": "Build Firecrawl Request",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Source Health Lookup": {
      "main": [
        [
          {
            "node": "Evaluate Dedup",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Read Reuse Route Rows": {
      "main": [
        [
          {
            "node": "Build Reuse Records",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build Reuse Records": {
      "main": [
        [
          {
            "node": "IF Reuse Records OK?",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "IF Reuse Records OK?": {
      "main": [
        [
          {
            "node": "Append Skipped Log (Duplicate)",
            "type": "main",
            "index": 0
          }
        ],
        [
          {
            "node": "Append Reuse Route Row",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Append Reuse Route Row": {
      "main": [
        [
          {
            "node": "Build Reuse Health Row",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build Reuse Health Row": {
      "main": [
        [
          {
            "node": "Append Reuse source_health",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Append Reuse source_health": {
      "main": [
        [
          {
            "node": "Loop Over Items",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Append live_source_runs": {
      "main": [
        [
          {
            "node": "Build agent_requests Row",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Append agent_requests": {
      "main": [
        [
          {
            "node": "Final Summary Output",
            "type": "main",
            "index": 0
          }
        ]
      ]
    }
  },
  "active": false,
  "settings": {
    "executionOrder": "v1"
  },
  "versionId": "b4000000-firecrawl-url-list-minibatch-v004-pts-contact-hardening-20260607"
}