AutomationFlowsWeb Scraping › Apify Search Candidate Discovery

Apify Search Candidate Discovery

05 - Apify Search Candidate Discovery. Uses httpRequest, googleSheets. Event-driven trigger; 15 nodes.

Event trigger★★★★☆ complexity15 nodesHTTP RequestGoogle Sheets
Web Scraping Trigger: Event Nodes: 15 Complexity: ★★★★☆ Added:

This workflow follows the Google Sheets → HTTP Request recipe pattern — see all workflows that pair these two integrations.

The workflow JSON

Copy or download the full n8n JSON below. Paste it into a new n8n workflow, add your credentials, activate. Full import guide →

Download .json
{
  "name": "05 - Apify Search Candidate Discovery",
  "nodes": [
    {
      "id": "n05-note-001",
      "name": "Overview Note RU",
      "type": "n8n-nodes-base.stickyNote",
      "typeVersion": 1,
      "position": [
        -380,
        -200
      ],
      "parameters": {
        "width": 520,
        "height": 520,
        "content": "## 05 \u2014 Apify: \u043f\u043e\u0438\u0441\u043a \u043a\u0430\u043d\u0434\u0438\u0434\u0430\u0442\u043e\u0432 URL (URL Supplier)\n\n\u041f\u043e\u0438\u0441\u043a\u043e\u0432\u044b\u0439 \u0437\u0430\u043f\u0440\u043e\u0441 \u2192 Apify Google Search Results Scraper \u2192 URL-\u043a\u0430\u043d\u0434\u0438\u0434\u0430\u0442\u044b \u2192 \u043f\u0440\u043e\u0432\u0435\u0440\u043a\u0430 url_registry \u2192\n\u0437\u0430\u043f\u0438\u0441\u044c \u0432 url_candidates + \u0441\u0442\u0440\u043e\u043a\u0430 \u0432 discovery_requests.\n\n\u041d\u0415 \u0430\u043d\u0430\u043b\u0438\u0437\u0438\u0440\u0443\u0435\u0442 \u0441\u0430\u0439\u0442\u044b. \u041d\u0415 \u0442\u0440\u0430\u0442\u0438\u0442 Firecrawl/Claude. \u041d\u0415 \u043e\u0431\u0440\u0430\u0431\u0430\u0442\u044b\u0432\u0430\u0435\u0442 \u0430\u0432\u0442\u043e\u043c\u0430\u0442\u0438\u0447\u0435\u0441\u043a\u0438.\n\u0422\u043e\u043b\u044c\u043a\u043e \u0420\u0423\u0427\u041d\u041e\u0419 \u0437\u0430\u043f\u0443\u0441\u043a, \u041e\u0414\u0418\u041d \u0437\u0430\u043f\u0440\u043e\u0441 \u0437\u0430 \u043f\u0440\u043e\u0433\u043e\u043d (v0.1), \u043c\u0430\u043a\u0441\u0438\u043c\u0443\u043c 10 \u043a\u0430\u043d\u0434\u0438\u0434\u0430\u0442\u043e\u0432.\n\n\u0427\u0442\u043e \u0434\u0435\u043b\u0430\u0435\u0442:\n1. Set Discovery Request \u2014 \u043e\u043f\u0435\u0440\u0430\u0442\u043e\u0440 \u0437\u0430\u0434\u0430\u0451\u0442 query / request_text (\u043f\u043e \u0443\u043c\u043e\u043b\u0447\u0430\u043d\u0438\u044e \u00ab\u0437\u0430\u0439\u043c \u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433 \u041f\u0422\u0421 \u041c\u043e\u0441\u043a\u0432\u0430\u00bb).\n2. Apify Google Search (sync) \u2192 \u043e\u0440\u0433\u0430\u043d\u0438\u0447\u0435\u0441\u043a\u0438\u0435 \u0440\u0435\u0437\u0443\u043b\u044c\u0442\u0430\u0442\u044b (url, title, snippet, rank).\n3. \u041d\u043e\u0440\u043c\u0430\u043b\u0438\u0437\u0443\u0435\u0442 URL (\u043a\u0430\u043a \u0432 Workflow 04 \u2192 \u043a\u043b\u044e\u0447 \u0441\u043e\u0432\u043f\u0430\u0434\u0430\u0435\u0442 \u0441 url_registry), \u0447\u0438\u0441\u0442\u0438\u0442 \u043c\u0443\u0441\u043e\u0440\u043d\u044b\u0435 \u0441\u0441\u044b\u043b\u043a\u0438.\n4. \u0427\u0438\u0442\u0430\u0435\u0442 url_registry (\u0438\u0441\u0442\u043e\u0447\u043d\u0438\u043a \u043f\u0440\u0430\u0432\u0434\u044b \u0434\u0435\u0434\u0443\u043f\u0430) \u0438 \u043f\u043e\u043c\u0435\u0447\u0430\u0435\u0442 dedup_status / registry_status.\n5. \u0414\u0435\u0442\u0435\u0440\u043c\u0438\u043d\u0438\u0440\u043e\u0432\u0430\u043d\u043d\u043e (\u0431\u0435\u0437 LLM) \u0441\u0447\u0438\u0442\u0430\u0435\u0442 confidence_score, region_hint, service_hint.\n6. \u041f\u0438\u0448\u0435\u0442 url_candidates (25 \u043a\u043e\u043b\u043e\u043d\u043e\u043a): approval_status=new \u0434\u043b\u044f \u0443\u043d\u0438\u043a\u0430\u043b\u044c\u043d\u044b\u0445, duplicate \u0434\u043b\u044f \u0434\u0443\u0431\u043b\u0435\u0439.\n7. \u041f\u0438\u0448\u0435\u0442 \u043e\u0434\u043d\u0443 \u0441\u0442\u0440\u043e\u043a\u0443 discovery_requests (18 \u043a\u043e\u043b\u043e\u043d\u043e\u043a): status=needs_review (\u0438\u043b\u0438 error).\n\n\u0414\u0430\u043b\u044c\u0448\u0435 \u2014 \u0420\u0423\u0427\u041d\u041e\u0415 \u043e\u0434\u043e\u0431\u0440\u0435\u043d\u0438\u0435: \u043e\u043f\u0435\u0440\u0430\u0442\u043e\u0440 \u0441\u0442\u0430\u0432\u0438\u0442 approval_status=approved \u0438 \u043f\u0435\u0440\u0435\u0434\u0430\u0451\u0442 \u22645 URL \u0432 Workflow 04.\nTelegram-\u0431\u043e\u0442 \u043f\u043e\u0437\u0436\u0435 \u0431\u0443\u0434\u0435\u0442 \u0438\u043d\u0442\u0435\u0440\u0444\u0435\u0439\u0441\u043e\u043c \u043d\u0430\u0434 discovery_requests + url_candidates + Workflow 04 (\u043d\u0435 \u0434\u0443\u0431\u043b\u0438\u0440\u0443\u044f \u043b\u043e\u0433\u0438\u043a\u0443).\n\n\u041b\u0438\u043c\u0438\u0442\u044b: \u0440\u0443\u0447\u043d\u043e\u0439 \u0442\u0440\u0438\u0433\u0433\u0435\u0440, \u0431\u0435\u0437 \u0440\u0430\u0441\u043f\u0438\u0441\u0430\u043d\u0438\u044f, \u0431\u0435\u0437 Firecrawl/Claude, \u0431\u0435\u0437 \u0437\u0430\u043f\u0438\u0441\u0438 \u0432 monitor_queue/results \u0438 \u0442.\u0434."
      }
    },
    {
      "id": "n05-note-002",
      "name": "Setup & Test RU",
      "type": "n8n-nodes-base.stickyNote",
      "typeVersion": 1,
      "position": [
        2080,
        -200
      ],
      "parameters": {
        "width": 460,
        "height": 380,
        "content": "## \u041d\u0430\u0441\u0442\u0440\u043e\u0439\u043a\u0430 \u0438 \u0442\u0435\u0441\u0442 (\u0432\u0440\u0443\u0447\u043d\u0443\u044e)\n\n1. \u041d\u0415 \u0430\u043a\u0442\u0438\u0432\u0438\u0440\u043e\u0432\u0430\u0442\u044c workflow (active=false).\n2. \u0421\u043e\u0437\u0434\u0430\u0442\u044c \u0432\u043a\u043b\u0430\u0434\u043a\u0438: discovery_requests (18 \u043a\u043e\u043b\u043e\u043d\u043e\u043a), url_candidates (25 \u043a\u043e\u043b\u043e\u043d\u043e\u043a). url_registry \u0443\u0436\u0435 \u0435\u0441\u0442\u044c (10 \u043a\u043e\u043b\u043e\u043d\u043e\u043a).\n3. \u0421\u043e\u0437\u0434\u0430\u0442\u044c \u043a\u0440\u0435\u0434\u0435\u043d\u0448\u043b Apify: 'Apify API - Marketing Scout' (Header Auth, Header Name=Authorization,\n   Value=Bearer <APIFY_API_TOKEN>, \u0434\u043e\u043c\u0435\u043d api.apify.com). \u0422\u043e\u043a\u0435\u043d \u0432\u0432\u043e\u0434\u0438\u0442\u044c \u0422\u041e\u041b\u042c\u041a\u041e \u0432 n8n, \u043d\u0435 \u0432 \u0444\u0430\u0439\u043b\u044b.\n4. \u041f\u041e\u0421\u041b\u0415 \u0418\u041c\u041f\u041e\u0420\u0422\u0410 \u043f\u0435\u0440\u0435\u043f\u0440\u0438\u0432\u044f\u0437\u0430\u0442\u044c \u043a\u0440\u0435\u0434\u0435\u043d\u0448\u043b\u044b:\n   - Apify Search API Request \u2192 Apify API - Marketing Scout\n   - Read url_registry / Append url_candidates / Append discovery_requests \u2192 Google Sheets - Marketing Scout Service Account\n5. \u041d\u0430 3 \u043d\u043e\u0434\u0430\u0445 Google Sheets \u0432\u0441\u0442\u0430\u0432\u0438\u0442\u044c \u0440\u0435\u0430\u043b\u044c\u043d\u044b\u0439 Spreadsheet ID (\u0437\u0430\u043c\u0435\u043d\u0438\u0442\u044c PASTE_SPREADSHEET_ID_HERE).\n6. \u0412 'Set Discovery Request' \u043f\u0440\u0438 \u0436\u0435\u043b\u0430\u043d\u0438\u0438 \u0438\u0437\u043c\u0435\u043d\u0438\u0442\u044c query / request_text.\n7. Execute Workflow \u043e\u0434\u0438\u043d \u0440\u0430\u0437.\n\n\u041e\u0436\u0438\u0434\u0430\u0435\u043c\u043e: 1 \u0441\u0442\u0440\u043e\u043a\u0430 discovery_requests (status=needs_review), \u0434\u043e 10 \u0441\u0442\u0440\u043e\u043a url_candidates,\n\u0434\u0443\u0431\u043b\u0438 \u043f\u043e\u043c\u0435\u0447\u0435\u043d\u044b approval_status=duplicate, \u043d\u0438\u043a\u0430\u043a\u0438\u0445 \u0437\u0430\u043f\u0438\u0441\u0435\u0439 \u0432 \u0431\u0438\u0437\u043d\u0435\u0441-\u0432\u043a\u043b\u0430\u0434\u043a\u0438, 0 Firecrawl/0 Claude."
      }
    },
    {
      "id": "n05-manual-01",
      "name": "Manual Start",
      "type": "n8n-nodes-base.manualTrigger",
      "typeVersion": 1,
      "position": [
        120,
        300
      ],
      "parameters": {}
    },
    {
      "id": "n05-set-req-01",
      "name": "Set Discovery Request",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        340,
        300
      ],
      "parameters": {
        "jsCode": "function pad(n){ return String(n).padStart(2,'0'); }\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nfunction moscowStamp(){var z=function(n){return String(n).padStart(2,'0');};var m=new Date(Date.now()+10800000);return m.getUTCFullYear()+z(m.getUTCMonth()+1)+z(m.getUTCDate())+'_'+z(m.getUTCHours())+z(m.getUTCMinutes())+z(m.getUTCSeconds());}\nconst now = moscowIsoNow();\nconst stamp = moscowStamp();\n\n// OPERATOR: edit query / request_text below before running. One query per run (v0.1).\nreturn [{ json: {\n  discovery_request_id: 'disc_'+stamp,\n  created_at: now,\n  requested_by: 'operator',\n  request_text: '\u043d\u0430\u0439\u0442\u0438 \u043a\u043e\u043d\u043a\u0443\u0440\u0435\u043d\u0442\u043e\u0432 \u043f\u043e \u0437\u0430\u0439\u043c\u0430\u043c \u043f\u043e\u0434 \u041f\u0422\u0421 \u0432 \u041c\u043e\u0441\u043a\u0432\u0435',\n  query: '\u0437\u0430\u0439\u043c \u043f\u043e\u0434 \u0437\u0430\u043b\u043e\u0433 \u041f\u0422\u0421 \u041c\u043e\u0441\u043a\u0432\u0430',\n  region: '\u041c\u043e\u0441\u043a\u0432\u0430',\n  service_focus: 'pts_loan',\n  requested_limit: 10,\n  source_mode: 'apify_search',\n  source_api: 'apify/google-search-scraper',\n  fixture_mode: false,            // \u00a74 set true for a dry run (no Apify call)\n  live_mode: true,\n  require_approval_token: true,   // \u00a74 paid Apify discovery requires an explicit approval token\n  approval_token: '',             // operator pastes the approved token for a live run (value never logged)\n  expected_approval_token: 'WF05_LIVE_APPROVED',\n  max_budget_usd: 0.20            // \u00a74 hard cost guard (>0 required)\n}}];"
      }
    },
    {
      "id": "n05-build-apify-01",
      "name": "Build Apify Search Request",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        560,
        300
      ],
      "parameters": {
        "jsCode": "const r = $json;\n// \u00a74 EXECUTABLE APPROVAL/BUDGET GATE (before any Apify call). An unapproved run THROWS here, so the\n// downstream HTTP node never executes => zero external calls. The token VALUE is never emitted.\n(function(){ function s(v){return v==null?'':String(v).trim();}\n  var fixture=(r.fixture_mode===true||s(r.fixture_mode).toLowerCase()==='true');\n  if(fixture)return; // fixture/dry runs never call Apify\n  var live=(r.live_mode!==false);\n  var needTok=(r.require_approval_token!==false);\n  var tokOk=(!needTok)||(s(r.approval_token)!==''&&s(r.approval_token)===s(r.expected_approval_token));\n  var budget=Number(r.max_budget_usd)||0;\n  if(!live)throw new Error('WF05 approval gate: live_mode=false \u2014 Apify discovery BLOCKED (zero external calls).');\n  if(!tokOk)throw new Error('WF05 approval gate: approval token missing/mismatch \u2014 Apify discovery BLOCKED (approval_token_used='+(s(r.approval_token)!==''?'provided_invalid':'no')+'; token value never logged).');\n  if(budget<=0)throw new Error('WF05 approval gate: max_budget_usd must be > 0 \u2014 Apify discovery BLOCKED.');\n})();\nconst query = String(r.query || '').trim();\nconst limit = Math.min(10, parseInt(r.requested_limit) || 10);\n\n// Apify Google Search Results Scraper input \u2014 minimal, low cost: one query, one page, organic only.\nconst apify_body = {\n  queries: query,\n  maxPagesPerQuery: 1,\n  resultsPerPage: limit,\n  countryCode: 'ru',\n  languageCode: 'ru',\n  includeUnfilteredResults: false,\n  saveHtml: false,\n  saveHtmlToKeyValueStore: false,\n  geminiSearch: { enableGemini: false },\n  perplexitySearch: { enablePerplexity: false, returnImages: false, returnRelatedQuestions: false },\n  chatGptSearch: { enableChatGpt: false },\n  copilotSearch: { enableCopilot: false },\n  maximumLeadsEnrichmentRecords: 0\n};\n\nreturn [{ json: { ...r, apify_body } }];"
      }
    },
    {
      "id": "n05-apify-http-01",
      "name": "Apify Search API Request",
      "type": "n8n-nodes-base.httpRequest",
      "typeVersion": 4.2,
      "position": [
        780,
        300
      ],
      "onError": "continueRegularOutput",
      "alwaysOutputData": true,
      "parameters": {
        "method": "POST",
        "url": "https://api.apify.com/v2/acts/apify~google-search-scraper/run-sync-get-dataset-items?format=json&clean=true",
        "authentication": "predefinedCredentialType",
        "nodeCredentialType": "httpHeaderAuth",
        "sendHeaders": true,
        "headerParameters": {
          "parameters": [
            {
              "name": "Content-Type",
              "value": "application/json"
            }
          ]
        },
        "sendBody": true,
        "contentType": "raw",
        "rawContentType": "application/json",
        "body": "={{ JSON.stringify($json.apify_body) }}",
        "options": {
          "timeout": 120000
        }
      },
      "credentials": {
        "httpHeaderAuth": {
          "name": "<your credential>"
        }
      }
    },
    {
      "id": "n05-normalize-01",
      "name": "Normalize Apify Results",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1000,
        300
      ],
      "parameters": {
        "jsCode": "const meta = $('Set Discovery Request').first().json;\nconst limit = Math.min(10, parseInt(meta.requested_limit) || 10);\nfunction cap(s, n){ return (s == null ? '' : String(s)).substring(0, n); }\n\n// Gather raw dataset items from whatever shape the HTTP node produced.\nlet datasetItems = [];\ntry {\n  const all = $input.all().map(i => i.json);\n  for (const it of all) {\n    if (it == null) continue;\n    if (Array.isArray(it)) datasetItems.push(...it);\n    else datasetItems.push(it);\n  }\n} catch (e) { datasetItems = []; }\n\n// Detect Apify error / empty response.\nlet ok = true; let errorPreview = '';\nconst first = datasetItems[0] || {};\nif (datasetItems.length === 0) { ok = false; errorPreview = 'Empty Apify response (no dataset items).'; }\nelse if (first && (first.error || (first.statusCode && first.statusCode >= 400) || first.success === false)) {\n  ok = false; errorPreview = cap('Apify error: ' + JSON.stringify(first.error || first.message || first.statusCode || first), 400);\n}\n\n// Collect organic-like results across dataset items.\nfunction collectOrganic(di){\n  const out = [];\n  if (Array.isArray(di.organicResults)) out.push(...di.organicResults);\n  if (Array.isArray(di.organic_results)) out.push(...di.organic_results);\n  if (Array.isArray(di.results)) out.push(...di.results);\n  if (out.length === 0 && (di.url || di.link)) out.push(di); // item is itself a result row\n  return out;\n}\nlet organic = [];\nfor (const di of datasetItems) { if (di && typeof di === 'object') organic.push(...collectOrganic(di)); }\n\n// Drop Google-internal / non-website / login junk URLs.\nfunction isJunk(u){\n  const l = String(u || '').toLowerCase();\n  if (!l) return true;\n  if (!/^https?:\\/\\//.test(l)) return true;\n  if (l.includes('google.') && (l.includes('/search') || l.includes('/url?') || l.includes('webcache') || l.includes('translate.google'))) return true;\n  if (l.includes('googleusercontent') || l.includes('gstatic.com')) return true;\n  if (l.includes('/maps') || l.startsWith('https://maps.')) return true;\n  if (l.includes('accounts.google') || l.includes('/signin') || l.includes('/login')) return true;\n  return false;\n}\n\n// URL normalizer \u2014 same rules as Workflow 04 (matches url_registry keys).\nfunction normalizeUrl(u){\n  u = String(u || '').trim();\n  if (!u) return '';\n  try {\n    const noFrag = u.split('#')[0];\n    const url = new URL(noFrag);\n    url.protocol = (url.protocol || '').toLowerCase();\n    url.hostname = (url.hostname || '').toLowerCase();\n    const drop = ['utm_source','utm_medium','utm_campaign','utm_term','utm_content','gclid','yclid','fbclid'];\n    for (const p of drop) url.searchParams.delete(p);\n    let path = url.pathname || '/';\n    if (path.length > 1 && path.endsWith('/')) path = path.slice(0, -1);\n    url.pathname = path;\n    return url.toString();\n  } catch (e) { return u; }\n}\n\nconst seen = new Set();\nconst candidates = [];\nlet rank = 0;\nfor (const r of organic) {\n  if (candidates.length >= limit) break;\n  const rawUrl = String((r && (r.url || r.link)) || '').trim();\n  if (!rawUrl || isJunk(rawUrl)) continue;\n  if (seen.has(rawUrl)) continue; // drop exact raw duplicates within the batch\n  seen.add(rawUrl);\n  rank++;\n  const norm = normalizeUrl(rawUrl);\n  let domain = '';\n  try { domain = new URL(norm).hostname; } catch (e) { domain = ''; }\n  if (!domain) { const m = String(rawUrl).match(/^https?:\\/\\/([^\\/?#]+)/i); domain = m ? m[1] : ''; }\n  domain = domain.toLowerCase(); if (domain.indexOf('www.') === 0) domain = domain.slice(4);\n  candidates.push({\n    candidate_url: rawUrl,\n    normalized_source_url: norm,\n    title: cap(r.title || r.titleText || '', 300),\n    snippet: cap(r.description || r.snippet || r.descriptionText || '', 600),\n    domain: domain,\n    rank: (r.position || r.rank || rank)\n  });\n}\nif (ok && candidates.length === 0) { ok = false; errorPreview = errorPreview || 'Apify returned no usable organic results.'; }\n\nreturn [{ json: {\n  ok: ok && candidates.length > 0,\n  error_preview: errorPreview,\n  candidate_count_raw: candidates.length,\n  candidates: candidates\n}}];"
      },
      "alwaysOutputData": true
    },
    {
      "id": "n05-read-reg-01",
      "name": "Read url_registry",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        1220,
        300
      ],
      "alwaysOutputData": true,
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "read",
        "documentId": {
          "__rl": true,
          "value": "PASTE_SPREADSHEET_ID_HERE",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "url_registry",
          "mode": "name"
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "n05-classify-01",
      "name": "Classify Candidates",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1440,
        300
      ],
      "parameters": {
        "jsCode": "const meta = $('Set Discovery Request').first().json;\nconst norm = $('Normalize Apify Results').first().json;\nconst candidates = Array.isArray(norm.candidates) ? norm.candidates : [];\n\n// Read url_registry once (source of truth for dedup). Do NOT scan business tabs.\nlet registry = [];\ntry { registry = $('Read url_registry').all().map(i => i.json).filter(Boolean); } catch (e) { registry = []; }\nconst regSet = new Set(registry.map(r => String(r.normalized_source_url || '').trim()).filter(Boolean));\n\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nfunction moscowStamp(){var z=function(n){return String(n).padStart(2,'0');};var m=new Date(Date.now()+10800000);return m.getUTCFullYear()+z(m.getUTCMonth()+1)+z(m.getUTCDate())+'_'+z(m.getUTCHours())+z(m.getUTCMinutes())+z(m.getUTCSeconds());}\nconst now = moscowIsoNow();\nconst stamp = moscowStamp();\nfunction clamp(n){ return Math.min(100, Math.max(1, Math.round(n))); }\nfunction low(v){ return String(v == null ? '' : v).toLowerCase(); }\n\nfunction extractDomain(u){\n  let h = '';\n  try { h = new URL(String(u||'').trim()).hostname; } catch (e) { h = ''; }\n  if (!h) { const m = String(u||'').match(/^https?:\\/\\/([^\\/?#]+)/i); h = m ? m[1] : ''; }\n  h = h.toLowerCase();\n  if (h.indexOf('www.') === 0) h = h.slice(4);\n  return h;\n}\n// Canonicalize a URL (S2-D6): lowercase scheme/host, strip www, drop fragment + tracking params, drop trailing slash.\nfunction canonicalize(u){\n  u = String(u||'').trim(); if (!u) return '';\n  try {\n    const url = new URL(u.split('#')[0]);\n    url.protocol = (url.protocol||'').toLowerCase();\n    url.hostname = (url.hostname||'').toLowerCase().replace(/^www\\./,'');\n    const drop = ['utm_source','utm_medium','utm_campaign','utm_term','utm_content','gclid','yclid','fbclid','_openstat'];\n    for (const p of drop) url.searchParams.delete(p);\n    let path = url.pathname || '/';\n    if (path.length > 1 && path.endsWith('/')) path = path.slice(0,-1);\n    url.pathname = path;\n    let out = url.toString();\n    if (out.endsWith('/') && (url.pathname === '/' || url.pathname === '')) out = out.replace(/\\/+$/,'');\n    return out;\n  } catch (e) { return u.split('#')[0].replace(/\\/+$/,''); }\n}\nfunction isRoot(u){\n  try { const p = new URL(String(u||'').trim()).pathname || '/'; return (p === '' || p === '/'); } catch (e) { return /^https?:\\/\\/[^\\/?#]+\\/?$/i.test(String(u||'').trim()); }\n}\n\nfunction regionHint(l){\n  if (l.includes('\u043c\u043e\u0441\u043a\u0432\u0430') || l.includes('\u043c\u043e\u0441\u043a\u043e\u0432\u0441\u043a') || l.includes('msk') || l.includes('moskva') || /(^|[^a-z\u0430-\u044f])\u043c\u043e([^a-z\u0430-\u044f]|$)/.test(l)) return '\u041c\u043e\u0441\u043a\u0432\u0430/\u041c\u041e';\n  return '';\n}\n// Canonical service detection (S2-D4): map evidence to canonical taxonomy services; broad brokerage wins.\nfunction canonService(s){s=low(s);var A={'credit_broker':'credit_brokerage','kreditnyy_broker':'credit_brokerage','generic_lending':'credit_brokerage','secured_auto_loan':'pts_loan','auto_collateral_loan':'pts_loan','zaym_pod_pts':'pts_loan','secured_real_estate_loan':'real_estate_secured_loan','zalog_nedvizhimosti':'real_estate_secured_loan','refinancing':'debt_refinancing','credit_refinancing':'debt_refinancing','ipoteka':'mortgage_brokerage','consumer_loan':'consumer_credit'};return A[s]||s;}\nconst SVC = [\n  ['credit_brokerage',['\u043a\u0440\u0435\u0434\u0438\u0442\u043d\u044b\u0439 \u0431\u0440\u043e\u043a\u0435\u0440','\u043f\u043e\u043c\u043e\u0449\u044c \u0432 \u043f\u043e\u043b\u0443\u0447\u0435\u043d\u0438','\u043f\u043e\u0434\u0431\u043e\u0440 \u0431\u0430\u043d\u043a','\u043f\u043e\u043c\u043e\u0449\u044c \u0441 \u043a\u0440\u0435\u0434\u0438\u0442','broker','\u0431\u0440\u043e\u043a\u0435\u0440']],\n  ['credit_after_refusals',['\u043f\u043e\u0441\u043b\u0435 \u043e\u0442\u043a\u0430\u0437\u043e\u0432','\u043e\u0442\u043a\u0430\u0437\u0430\u043b\u0438','\u043f\u043b\u043e\u0445\u0430\u044f \u043a\u0440\u0435\u0434\u0438\u0442\u043d','\u043f\u0440\u043e\u0441\u0440\u043e\u0447\u043a']],\n  ['pts_loan',['\u043f\u0442\u0441','pts','\u0437\u0430\u043b\u043e\u0433 \u0430\u0432\u0442\u043e','\u043f\u043e\u0434 \u0430\u0432\u0442\u043e','pod-zalog-pts','zalog-pts']],\n  ['real_estate_secured_loan',['\u043d\u0435\u0434\u0432\u0438\u0436','\u043a\u0432\u0430\u0440\u0442\u0438\u0440','nedvizh','\u0437\u0430\u043b\u043e\u0433 \u043d\u0435\u0434\u0432\u0438\u0436']],\n  ['mortgage_brokerage',['\u0438\u043f\u043e\u0442\u0435\u043a','\u0438\u043f\u043e\u0442\u0435\u0447\u043d','ipotek']],\n  ['debt_refinancing',['\u0440\u0435\u0444\u0438\u043d\u0430\u043d\u0441','refinanc']],\n  ['business_credit',['\u0434\u043b\u044f \u0431\u0438\u0437\u043d\u0435\u0441\u0430','\u043e\u0431\u043e\u0440\u043e\u0442\u043d','\u0442\u0435\u043d\u0434\u0435\u0440\u043d','\u0438\u043f \u0438 \u043e\u043e\u043e']],\n  ['bank_guarantee',['\u0431\u0430\u043d\u043a\u043e\u0432\u0441\u043a \u0433\u0430\u0440\u0430\u043d\u0442','\u0433\u0430\u0440\u0430\u043d\u0442\u0438']],\n  ['auto_credit',['\u0430\u0432\u0442\u043e\u043a\u0440\u0435\u0434\u0438\u0442']],\n  ['consumer_credit',['\u043f\u043e\u0442\u0440\u0435\u0431\u0438\u0442\u0435\u043b\u044c\u0441\u043a','\u043f\u043e\u0442\u0440\u0435\u0431 \u043a\u0440\u0435\u0434\u0438\u0442']]\n];\nfunction detectServices(blob){ const out=[]; for (const p of SVC){ if (p[1].some(t=>blob.indexOf(t)>=0)) out.push(p[0]); } return out; }\nfunction looksLender(l){\n  return l.includes('\u043b\u043e\u043c\u0431\u0430\u0440\u0434') || l.includes('lombard') || l.includes('autolombard') || l.includes('\u0437\u0430\u0439\u043c') || l.includes('\u0437\u0430\u0451\u043c') || l.includes('\u0437\u0430\u0438\u043c') || l.includes('\u043a\u0440\u0435\u0434\u0438\u0442') || l.includes('\u0444\u0438\u043d\u0430\u043d\u0441') || l.includes('finance') || l.includes('credit') || l.includes('loan');\n}\n\n// ---- request-level scope/service representation (S2-D1) ----\nconst queryLow = low(meta.query);\nconst queryServices = detectServices(queryLow);\nconst focusCanon = canonService(meta.service_focus);\nconst broadBrokerage = /\u0431\u0440\u043e\u043a\u0435\u0440|\u043f\u043e\u043c\u043e\u0449\u044c \u0432 \u043f\u043e\u043b\u0443\u0447\u0435\u043d\u0438|\u043f\u043e\u0434\u0431\u043e\u0440 \u0431\u0430\u043d\u043a/.test(queryLow);\nlet requestServicePrimary = broadBrokerage ? 'credit_brokerage' : (queryServices[0] || (focusCanon && focusCanon !== 'unknown' ? focusCanon : 'unknown'));\nconst requestServiceSecondaryArr = [];\nfor (const s of queryServices) if (s !== requestServicePrimary && requestServiceSecondaryArr.indexOf(s) < 0) requestServiceSecondaryArr.push(s);\nif (focusCanon && focusCanon !== 'unknown' && focusCanon !== requestServicePrimary && requestServiceSecondaryArr.indexOf(focusCanon) < 0) requestServiceSecondaryArr.push(focusCanon);\nconst requested_search_scope = broadBrokerage ? 'broad_credit_brokerage' : ('narrow:' + requestServicePrimary);\nconst query_terms = String(meta.query||'').toLowerCase().split(/[^a-z\u0430-\u044f\u04510-9]+/).filter(w => w && w.length >= 3).slice(0, 12);\n\n// ---- candidate classification (S2-D2): regulator / publisher / source_candidate / direct / indirect / irrelevant ----\nconst REGULATORS = ['cbr.ru','gov.ru','nalog.ru','nalog.gov.ru','consultant.ru','garant.ru','gosuslugi.ru','fssp.gov.ru','sudrf.ru','minfin.ru','asv.org.ru'];\nconst AGGREGATORS = ['banki.ru','vbr.ru','finuslugi.ru','sravni.ru','bankiros','myfin','vsezaimy','vse-zaimy','zaim.com','outbank','frbank','brobank'];\nconst DIRECTORIES = ['2gis.','zoon.','yell.ru','flamp.','orgpage','rusprofile','spr.ru','yandex.ru/maps','/maps'];\nconst MEDIA = ['kp.ru','rbc.ru','ria.ru','lenta.','vc.ru','dzen.ru','journal.tinkoff','pikabu','habr.','iz.ru','aif.ru','gazeta.ru'];\nconst MARKETPLACES = ['avito.','youla.','yula.','ozon.','wildberries','market.yandex','drom.ru','auto.ru'];\nconst SOCIALS = ['vk.com','ok.ru','instagram.','t.me','telegram','facebook.','youtube.','tiktok.','twitter.','x.com','dzen.ru/id'];\nfunction candidateType(domain, base){\n  const dl = domain.toLowerCase();\n  if (REGULATORS.some(x => dl === x || dl.endsWith('.' + x) || dl.includes(x))) return 'regulator';\n  if (SOCIALS.some(x => dl.includes(x))) return 'social';\n  if (MARKETPLACES.some(x => dl.includes(x))) return 'marketplace';\n  if (DIRECTORIES.some(x => dl.includes(x) || base.includes(x))) return 'directory';\n  if (AGGREGATORS.some(x => dl.includes(x))) return 'aggregator';\n  if (MEDIA.some(x => dl.includes(x))) return 'publisher_article';\n  if (/\\/(news|article|articles|blog|stati|statya|reviews|rating|reyting|top-?\\d|luchshi|obzor)(\\/|$)/.test(base)) return 'publisher_article';\n  if (looksLender(base) || dl.includes('lombard') || dl.includes('zalog') || dl.includes('credit') || dl.includes('cash') || dl.includes('finans') || dl.includes('zaim')) return 'direct_competitor';\n  return 'unknown';\n}\n// competitor_class: how this candidate may be USED. Regulators/publishers are NEVER competitor entities.\nfunction competitorClass(ct){\n  if (ct === 'direct_competitor') return 'direct';\n  if (ct === 'aggregator') return 'indirect';\n  if (ct === 'regulator') return 'regulator';\n  if (ct === 'publisher_article') return 'publisher';\n  if (ct === 'directory' || ct === 'marketplace' || ct === 'social') return 'source_candidate';\n  return 'irrelevant';\n}\n\nconst wantsMarket = /\u0430\u0432\u0438\u0442\u043e|\u043c\u0430\u0440\u043a\u0435\u0442\u043f\u043b\u0435\u0439\u0441|\u043e\u0431\u044a\u044f\u0432\u043b\u0435\u043d|avito|youla|\u044e\u043b\u0430/.test(queryLow);\nconst batchSeen = new Set();\nconst rows = [];\nlet unique = 0, dups = 0, estFc = 0, estClaude = 0;\nconst classCounts = {};\n\nfor (let i = 0; i < candidates.length; i++) {\n  const c = candidates[i];\n  const canonical_url = canonicalize(c.normalized_source_url || c.candidate_url);\n  const key = String(c.normalized_source_url || '').trim() || canonical_url;\n  const domain = extractDomain(canonical_url || c.candidate_url);\n  const base = ((c.title||'') + ' ' + (c.snippet||'') + ' ' + domain + ' ' + (c.candidate_url||'')).toLowerCase();\n  const hay = base + ' ' + queryLow;\n\n  const inRegistry = key !== '' && regSet.has(key);\n  let dedup_status = 'unique';\n  if (inRegistry) dedup_status = 'duplicate_in_registry';\n  else if (key !== '' && batchSeen.has(key)) dedup_status = 'duplicate_in_batch';\n  if (!inRegistry && key !== '') batchSeen.add(key);\n  const registry_status = inRegistry ? 'in_registry' : 'not_in_registry';\n  const isDup = (dedup_status !== 'unique');\n  const approval_status = isDup ? 'duplicate' : 'new';\n\n  const candidate_type = candidateType(domain, base);\n  const competitor_class = competitorClass(candidate_type);\n  const is_competitor_entity = (competitor_class === 'direct' || competitor_class === 'indirect');\n  classCounts[competitor_class] = (classCounts[competitor_class] || 0) + 1;\n\n  // per-candidate canonical service hint(s) from candidate evidence (not the query alone)\n  const candServices = detectServices(base);\n  const service_primary = candServices[0] || (is_competitor_entity ? requestServicePrimary : 'unknown');\n  const service_secondary = candServices.slice(1, 3).join(', ');\n\n  // Confidence (S2-D3): domain/title/snippet/path/service + class. Regulators/publishers never look like a direct competitor.\n  let s = 10;\n  if (candidate_type === 'direct_competitor') s += 30;\n  if (base.includes('\u0437\u0430\u043b\u043e\u0433')) s += 15;\n  if (base.includes('\u043f\u0442\u0441') || base.includes('pts')) s += 15;\n  if (base.includes('\u0430\u0432\u0442\u043e') || base.includes('\u0430\u0432\u0442\u043e\u043c\u043e\u0431\u0438\u043b')) s += 10;\n  if (regionHint(hay)) s += 12;\n  if (base.includes('\u043a\u0440\u0435\u0434\u0438\u0442') || base.includes('\u0437\u0430\u0439\u043c') || base.includes('\u0437\u0430\u0451\u043c')) s += 8;\n  if (base.includes('\u0441\u0442\u0430\u0432\u043a\u0430') || base.includes('\u0441\u0443\u043c\u043c\u0430') || base.includes('\u043e\u0434\u043e\u0431\u0440\u0435\u043d')) s += 8;\n  if (service_primary !== 'unknown') s += 8;\n  if (looksLender(base)) s += 8;\n  if (dedup_status === 'duplicate_in_registry') s -= 50;\n  if (candidate_type === 'regulator') s -= 45;\n  if (candidate_type === 'publisher_article') s -= 30;\n  if (candidate_type === 'directory') s -= 30;\n  if (candidate_type === 'aggregator') s -= 20;\n  if (candidate_type === 'marketplace' && !wantsMarket) s -= 20;\n  if (candidate_type === 'social') s -= 28;\n  const confidence_score = clamp(s);\n\n  const est_fc = (approval_status === 'new') ? 1 : 0;\n  const est_cl = (approval_status === 'new') ? 0.02 : 0;\n  if (approval_status === 'new') { unique++; estFc += est_fc; estClaude += est_cl; } else { dups++; }\n\n  let note;\n  if (isDup) note = 'Duplicate; do not process unless force_reprocess later. Requires human approval before Workflow 04.';\n  else if (candidate_type === 'regulator') note = 'Regulator/government source (e.g. cbr.ru) \u2014 NOT a competitor; useful as a monitoring/source candidate only. Do not write as a competitor entity.';\n  else if (candidate_type === 'publisher_article') note = 'Publisher/article \u2014 NOT a competitor; may be a source candidate for audience/market signals. Do not write as a competitor entity.';\n  else if (!is_competitor_entity) note = 'Source candidate (' + candidate_type + '); review manually. Not a competitor entity.';\n  else if (candidate_type !== 'direct_competitor') note = 'Indirect competitor (' + candidate_type + '); review before Workflow 04.';\n  else note = 'Direct competitor; requires human approval before Workflow 04.';\n\n  rows.push({\n    candidate_id: 'cand_' + stamp + '_' + (i + 1),\n    discovery_request_id: meta.discovery_request_id,\n    created_at: now,\n    requested_by: meta.requested_by || 'operator',\n    requested_limit: meta.requested_limit || 10,\n    query: meta.query || '',\n    requested_search_scope: requested_search_scope,\n    source: 'apify_search',\n    candidate_url: c.candidate_url || '',\n    normalized_source_url: key,\n    canonical_url: canonical_url,\n    is_root: isRoot(c.candidate_url || canonical_url),\n    domain: domain,\n    candidate_type: candidate_type,\n    competitor_class: competitor_class,\n    is_competitor_entity: is_competitor_entity,\n    title: c.title || '',\n    snippet: c.snippet || '',\n    rank: c.rank || (i + 1),\n    region_hint: regionHint(hay),\n    service_hint: service_primary,\n    service_primary: service_primary,\n    service_secondary: service_secondary,\n    confidence_score: confidence_score,\n    dedup_status: dedup_status,\n    registry_status: registry_status,\n    approval_status: approval_status,\n    approved_by: '',\n    approved_at: '',\n    rejection_reason: '',\n    estimated_firecrawl_credits: est_fc,\n    estimated_claude_cost_usd: est_cl,\n    notes: note\n  });\n}\n\nreturn [{ json: {\n  meta: meta,\n  ok: (norm.ok === true) && rows.length > 0,\n  error_preview: norm.error_preview || '',\n  requested_search_scope: requested_search_scope,\n  service_primary: requestServicePrimary,\n  service_secondary: requestServiceSecondaryArr.join(', '),\n  query_terms: query_terms,\n  candidateRows: rows,\n  candidate_count: rows.length,\n  unique_candidate_count: unique,\n  duplicate_count: dups,\n  class_counts: classCounts,\n  direct_competitor_count: classCounts.direct || 0,\n  regulator_count: classCounts.regulator || 0,\n  publisher_count: classCounts.publisher || 0,\n  source_candidate_count: classCounts.source_candidate || 0,\n  estimated_firecrawl_credits: estFc,\n  estimated_claude_cost_usd: Math.round(estClaude * 1000) / 1000,\n  estimated_apify_cost_usd: null,\n  actual_apify_cost_usd: null,\n  apify_cost_status: 'unknown'\n}}];"
      },
      "alwaysOutputData": true
    },
    {
      "id": "n05-expand-01",
      "name": "Expand Candidate Rows",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1660,
        300
      ],
      "parameters": {
        "jsCode": "const c = $json;\nconst rows = Array.isArray(c.candidateRows) ? c.candidateRows : [];\n// Fan out 0..10 candidate rows for Append url_candidates. 0 \u2192 nothing appended.\nreturn rows.map(r => ({ json: r }));"
      }
    },
    {
      "id": "n05-append-cand-01",
      "name": "Append url_candidates",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        1880,
        300
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "PASTE_SPREADSHEET_ID_HERE",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "url_candidates",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "n05-summary-01",
      "name": "Build Discovery Request Summary",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1660,
        520
      ],
      "parameters": {
        "jsCode": "const c = $('Classify Candidates').first().json;\nconst meta = c.meta || $('Set Discovery Request').first().json;\nfunction moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\nconst now = moscowIsoNow();\nconst count = c.candidate_count || 0;\nconst status = (c.ok === true && count > 0) ? 'needs_review' : 'error';\n\nlet notes = (status === 'needs_review')\n  ? ('Discovery via Apify Google Search. ' + count + ' candidates (' + (c.unique_candidate_count||0) + ' new, ' + (c.duplicate_count||0) + ' duplicate). Requires human approval before Workflow 04.')\n  : ('Apify discovery returned no usable candidates. ' + (c.error_preview || ''));\nnotes = String(notes).substring(0, 500);\n\nreturn [{ json: {\n  discovery_request_id: meta.discovery_request_id,\n  created_at: meta.created_at || now,\n  requested_by: meta.requested_by || 'operator',\n  request_text: meta.request_text || '',\n  query: meta.query || '',\n  region: meta.region || '',\n  service_focus: meta.service_focus || '',\n  requested_limit: meta.requested_limit || 10,\n  source_mode: 'apify_search',\n  source_api: 'apify/google-search-scraper',\n  status: status,\n  candidate_count: count,\n  unique_candidate_count: c.unique_candidate_count || 0,\n  duplicate_count: c.duplicate_count || 0,\n  approved_count: 0,\n  estimated_firecrawl_credits: c.estimated_firecrawl_credits || 0,\n  requested_search_scope: c.requested_search_scope || '',\n  service_primary: c.service_primary || '',\n  service_secondary: c.service_secondary || '',\n  query_terms: Array.isArray(c.query_terms) ? c.query_terms.join(', ') : '',\n  direct_competitor_count: c.direct_competitor_count || 0,\n  regulator_count: c.regulator_count || 0,\n  publisher_count: c.publisher_count || 0,\n  source_candidate_count: c.source_candidate_count || 0,\n  estimated_firecrawl_credits: c.estimated_firecrawl_credits || 0,\n  estimated_claude_cost_usd: c.estimated_claude_cost_usd || 0,\n  estimated_apify_cost_usd: (c.estimated_apify_cost_usd == null ? '' : c.estimated_apify_cost_usd),\n  actual_apify_cost_usd: (c.actual_apify_cost_usd == null ? '' : c.actual_apify_cost_usd),\n  apify_cost_status: c.apify_cost_status || 'unknown',\n  notes: notes\n}}];"
      }
    },
    {
      "id": "n05-append-req-01",
      "name": "Append discovery_requests",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        1880,
        520
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "PASTE_SPREADSHEET_ID_HERE",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "discovery_requests",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    },
    {
      "id": "wf05-build-runlog",
      "name": "Build live_source_runs Row",
      "type": "n8n-nodes-base.code",
      "typeVersion": 2,
      "position": [
        1600,
        820
      ],
      "parameters": {
        "jsCode": "function moscowIsoNow(){var m=new Date(Date.now()+10800000);return m.toISOString().replace('Z','+03:00');}\n// WF05 live_source_runs ledger (23 cols, DEC-126). Apify search discovery run. No new external calls here.\nfunction str(v){return v==null?'':String(v).trim();}\nconst s=$('Build Discovery Request Summary').first().json;\nconst cfg=$('Set Discovery Request').first().json||{};\nconst cnt=Number(s.candidate_count)||0;\nconst directCompetitors=Number(s.direct_competitor_count)||0;\nvar _needTok=(cfg.require_approval_token!==false);\nvar _tokOk=(str(cfg.approval_token)!==''&&str(cfg.approval_token)===str(cfg.expected_approval_token));\nvar approvalTokenUsed=_needTok?(_tokOk?'yes':'no'):'not_required';\nconst st=(str(s.status)==='needs_review')?'completed':(cnt>0?'partial':'failed');\nreturn [{ json: {\n  run_id:str(s.discovery_request_id)||('disc_'+Date.now()),\n  created_at:moscowIsoNow(),\n  workflow:'WF05 - Apify Search Candidate Discovery',\n  source_family:'web_competitor',\n  platform:'apify_search',\n  mode:'live',\n  approval_token_used:approvalTokenUsed,\n  source_allowlist:str(s.query||cfg.query),\n  max_items:Number(s.requested_limit)||Number(cfg.requested_limit)||10,\n  items_received:cnt,\n  items_relevant:directCompetitors,\n  items_written_raw:0,\n  items_unique:Number(s.unique_candidate_count)||0,\n  items_duplicate:Number(s.duplicate_count)||0,\n  hard_skipped:0,\n  external_calls:1,\n  apify_cost_status:'unknown',\n  actual_source_cost_usd:null,\n  estimated_source_cost_usd:'',\n  llm_calls:0,\n  estimated_llm_cost_usd:0,\n  status:st,\n  error_summary:(st==='failed'?str(s.notes):''),\n  operator_next_action:'Approve direct_competitor candidates in url_candidates, then run WF06.',\n  notes:'Apify google-search discovery; cost not recovered (apify_search). Writes url_candidates (not raw_market_records).'\n}}];"
      }
    },
    {
      "id": "wf05-app-runlog",
      "name": "Append live_source_runs",
      "type": "n8n-nodes-base.googleSheets",
      "typeVersion": 4,
      "position": [
        1820,
        820
      ],
      "parameters": {
        "authentication": "serviceAccount",
        "operation": "append",
        "documentId": {
          "__rl": true,
          "value": "PASTE_SPREADSHEET_ID_HERE",
          "mode": "id"
        },
        "sheetName": {
          "__rl": true,
          "value": "live_source_runs",
          "mode": "name"
        },
        "columns": {
          "mappingMode": "autoMapInputData",
          "value": {},
          "matchingColumns": [],
          "schema": []
        },
        "options": {}
      },
      "credentials": {
        "googleApi": {
          "name": "<your credential>"
        }
      },
      "retryOnFail": true,
      "maxTries": 3,
      "waitBetweenTries": 5000
    }
  ],
  "connections": {
    "Manual Start": {
      "main": [
        [
          {
            "node": "Set Discovery Request",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Set Discovery Request": {
      "main": [
        [
          {
            "node": "Build Apify Search Request",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build Apify Search Request": {
      "main": [
        [
          {
            "node": "Apify Search API Request",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Apify Search API Request": {
      "main": [
        [
          {
            "node": "Normalize Apify Results",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Normalize Apify Results": {
      "main": [
        [
          {
            "node": "Read url_registry",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Read url_registry": {
      "main": [
        [
          {
            "node": "Classify Candidates",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Classify Candidates": {
      "main": [
        [
          {
            "node": "Expand Candidate Rows",
            "type": "main",
            "index": 0
          },
          {
            "node": "Build Discovery Request Summary",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Expand Candidate Rows": {
      "main": [
        [
          {
            "node": "Append url_candidates",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build Discovery Request Summary": {
      "main": [
        [
          {
            "node": "Append discovery_requests",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Append discovery_requests": {
      "main": [
        [
          {
            "node": "Build live_source_runs Row",
            "type": "main",
            "index": 0
          }
        ]
      ]
    },
    "Build live_source_runs Row": {
      "main": [
        [
          {
            "node": "Append live_source_runs",
            "type": "main",
            "index": 0
          }
        ]
      ]
    }
  },
  "active": false,
  "settings": {
    "executionOrder": "v1"
  },
  "versionId": "05-apify-search-candidate-discovery-v0-1"
}

Credentials you'll need

Each integration node will prompt for credentials when you import. We strip credential IDs before publishing — you'll add your own.

Pro

For the full experience including quality scoring and batch install features for each workflow upgrade to Pro

About this workflow

05 - Apify Search Candidate Discovery. Uses httpRequest, googleSheets. Event-driven trigger; 15 nodes.

Source: https://github.com/CodeVinci8/vinci-ai-pilot/blob/main/n8n/workflows/05_apify_search_candidate_discovery.json — original creator credit. Request a take-down →

More Web Scraping workflows → · Browse all categories →

Related workflows

Workflows that share integrations, category, or trigger type with this one. All free to copy and import.

Web Scraping

04 - Firecrawl URL List Mini-Batch to Resilient Analyzer. Uses googleSheets, httpRequest, executeWorkflowTrigger. Event-driven trigger; 42 nodes.

Google Sheets, HTTP Request, Execute Workflow Trigger
Web Scraping

Automate LinkedIn lead generation by scraping comments from targeted posts and enriching profiles with detailed data

Form Trigger, HTTP Request, Google Sheets
Web Scraping

This automated n8n workflow scrapes job listings from Upwork using Apify, processes and cleans the data, and generates daily email reports with job summaries. The system uses Google Sheets for data st

Google Sheets, HTTP Request, Gmail
Web Scraping

Transform LinkedIn profile URLs into comprehensive enriched lead profiles, quickly and automatically.

HTTP Request, Google Sheets
Web Scraping

Transform any website into a structured knowledge repository with this intelligent crawler that extracts hyperlinks from the homepage, intelligently filters images and content pages, and aggregates fu

HTTP Request, Google Sheets