Commit Graph

  • 63787bc504 fix(scrapeURL/fire-engine): wait longer if timeout is not specified main Gergő Móricz 2024-11-15 20:25:16 +0100
  • 4cddcd5206 fix(scrapeURL/fire-engine): timeout-less scrape support (initial) Gergő Móricz 2024-11-15 20:15:25 +0100
  • 350d00d27a fix(crawler): treat XML files as sitemaps (temporarily) Gergő Móricz 2024-11-15 20:09:20 +0100
  • ca2e33db0a fix(log_job): add force option to retry on supabase failure Gergő Móricz 2024-11-15 19:55:23 +0100
  • 7b02c45dd0 fix(v1/types): better timeout primitives Gergő Móricz 2024-11-15 19:35:54 +0100
  • b1302d6221
    Merge 6f45ab6691 into c95a4a26c9 Rui Rua 2024-11-15 12:59:36 -0500
  • c95a4a26c9 fix(v1/batch/scrape): raise default timeout Gergő Móricz 2024-11-15 18:58:03 +0100
  • d9c4bb5aaf
    Merge f6e6f2ef9f into 3a342bfbf0 Nicolas 2024-11-15 22:59:49 +0530
  • 3f2914b97f
    Merge 9298a05045 into 3a342bfbf0 Gergő Móricz 2024-11-15 22:54:35 +0530
  • 3a342bfbf0 fix(scrapeURL/playwright): JSON body fix Móricz Gergő 2024-11-15 15:18:40 +0100
  • 18352bc153
    Merge 5453539fb4 into 3c1b1909f8 dependabot[bot] 2024-11-15 15:34:43 +0800
  • 3c1b1909f8 Update map.ts Nicolas 2024-11-14 17:52:15 -0500
  • 9519897102 Merge branch 'nsc/sitemap-only' Nicolas 2024-11-14 17:44:39 -0500
  • 7f084c6c43 Nick: nsc/sitemap-only Nicolas 2024-11-14 17:44:32 -0500
  • e8bd089c8a
    Merge pull request #901 from mendableai/nsc/sitemap-only Nicolas 2024-11-14 17:32:37 -0500
  • 3fcdf57d2f Update fireEngine.ts Nicolas 2024-11-14 17:31:30 -0500
  • d62f12c9d9 Nick: moved away from axios Nicolas 2024-11-14 17:31:23 -0500
  • f155449458 Nick: sitemap only Nicolas 2024-11-14 17:29:53 -0500
  • 431e64e752 fix(batch/scrape/webhook): add batch_scrape.started Móricz Gergő 2024-11-14 22:39:41 +0100
  • 7bca4486b4 Update package.json Nicolas 2024-11-14 16:37:53 -0500
  • df05124ef5 feat(v1/batch/scrape): webhooks Móricz Gergő 2024-11-14 22:36:28 +0100
  • 91f4fd815f Update extract.ts nsc/new-extract Nicolas 2024-11-14 15:41:42 -0500
  • 22848af5ae Nick: Nicolas 2024-11-14 15:34:02 -0500
  • ebe9de2ac5 Nick: Nicolas 2024-11-14 15:26:15 -0500
  • 5056dcd8e9 Update index.test.ts Nicolas 2024-11-14 15:06:22 -0500
  • be5e6da5ff tests rafaelmmiller 2024-11-14 17:03:54 -0300
  • 45091430ab Merge branch 'nsc/new-extract' of https://github.com/mendableai/firecrawl into nsc/new-extract rafaelmmiller 2024-11-14 17:03:15 -0300
  • 796cd0746d Update extract.ts Nicolas 2024-11-14 15:03:06 -0500
  • 1b5f6a0959 Update extract.ts Nicolas 2024-11-14 14:59:34 -0500
  • d6749c211d Nick: refactor and /* glob pattern support Nicolas 2024-11-14 14:57:38 -0500
  • 41b45a844b sdk allowexternallinks rafaelmmiller 2024-11-14 15:56:12 -0300
  • 3d6d650f0b Merge branch 'nsc/new-extract' of https://github.com/mendableai/firecrawl into nsc/new-extract rafaelmmiller 2024-11-14 15:51:32 -0300
  • 80d6cb16fb sdks wip rafaelmmiller 2024-11-14 15:51:27 -0300
  • 359c30fbda fix(cache): don't cache on failure error code Gergő Móricz 2024-11-14 19:49:34 +0100
  • 49ff37afb4 feat: cache Gergő Móricz 2024-11-14 19:47:12 +0100
  • 86a78a03cb fix(sitemap): scrape with tlsclient Gergő Móricz 2024-11-11 19:44:32 +0100
  • 62c8b63b84 Create README.md Eric Ciarla 2024-11-14 11:55:00 -0500
  • 9298a05045 feat: turn into API mog/doctor Móricz Gergő 2024-11-14 12:24:53 +0100
  • 5519f077aa fix(scrapeURL): adjust error message for clarity Móricz Gergő 2024-11-14 10:13:48 +0100
  • faf11acf82 doctor first iteration Móricz Gergő 2024-11-14 10:12:49 +0100
  • 0a1c99074f fix(html-to-markdown): make error reporting less intrusive Móricz Gergő 2024-11-14 08:58:00 +0100
  • bd928b1512 Nick: changed email from hello to help Nicolas 2024-11-13 20:27:20 -0500
  • a1c018fdb0 Merge branch 'main' into nsc/new-extract Nicolas 2024-11-13 17:14:43 -0500
  • 904c904971 wip rafaelmmiller 2024-11-13 18:06:20 -0300
  • 0310cd2afa fix(crawl): redirect rebase Gergő Móricz 2024-11-13 21:38:44 +0100
  • 0d1c4e4e09 Update package.json Nicolas 2024-11-13 13:54:22 -0500
  • 32be2cf786
    feat(v1/webhook): complex webhook object w/ headers (#899) Gergő Móricz 2024-11-13 19:36:44 +0100
  • 4470cdf731 feat(js-sdk/crawl): add complex webhook support mog/complex-webhook Gergő Móricz 2024-11-13 19:36:28 +0100
  • 1c1ac7ced3 feat(v1/webhook): complex webhook object w/ headers Gergő Móricz 2024-11-13 19:04:08 +0100
  • ea1302960f
    Merge pull request #895 from mendableai/nsc/redlock-email Nicolas 2024-11-13 12:45:55 -0500
  • 25f32000db mvp done? rafaelmmiller 2024-11-13 13:05:29 -0300
  • a175c1513a wip rafaelmmiller 2024-11-13 08:09:51 -0300
  • 775f04610a
    Merge d301c1bf0f into 5ce4aaf0ec Meshari 2024-11-13 12:01:00 +0300
  • 1a636b4e59 Update email_notification.ts nsc/redlock-email Nicolas 2024-11-12 20:09:01 -0500
  • efa89b5a5d
    Merge e7d73f1c0d into 5ce4aaf0ec dependabot[bot] 2024-11-12 23:52:03 +0000
  • 5ce4aaf0ec fix(crawl): initialURL setting is unnecessary Gergő Móricz 2024-11-12 23:35:07 +0100
  • 5453539fb4
    apps/api(deps): bump the prod-deps group across 1 directory with 40 updates dependabot/npm_and_yarn/apps/api/prod-deps-65027e8b4e dependabot[bot] 2024-11-12 21:54:15 +0000
  • 93ac20f930 fix(queue-worker): do not kill crawl on one-page error Gergő Móricz 2024-11-12 22:53:29 +0100
  • e7d73f1c0d
    apps/api(deps-dev): bump the dev-deps group across 1 directory with 12 updates dependabot/npm_and_yarn/apps/api/dev-deps-0f7ee87ebe dependabot[bot] 2024-11-12 21:50:16 +0000
  • 16e850288c fix(scrapeURL/pdf,docx): ignore SSL when downloading PDF Gergő Móricz 2024-11-12 22:46:58 +0100
  • 807703d94c wip rafaelmmiller 2024-11-12 18:44:14 -0300
  • 7081beff1f fix(scrapeURL/pdf): retry Gergő Móricz 2024-11-12 22:26:36 +0100
  • 9ace2ad071 fix(scrapeURL/pdf): fix llamaparse upload Gergő Móricz 2024-11-12 20:55:14 +0100
  • 687ea69621 fix(requests.http): default to localhost baseUrl Gergő Móricz 2024-11-12 19:59:09 +0100
  • 3a5eee6e3f feat: improve requests.http using format features Gergő Móricz 2024-11-12 19:58:07 +0100
  • f2ecf0cc36 fix(v0): crawl timeout errors Gergő Móricz 2024-11-12 19:46:00 +0100
  • 464b41a5d2 Merge branch 'main' into nsc/new-extract Nicolas 2024-11-12 12:24:47 -0500
  • a23364e5da Update extract.ts Nicolas 2024-11-12 12:23:44 -0500
  • a4f15260a7 Nick: Nicolas 2024-11-12 12:23:24 -0500
  • fbabc779f5
    fix(crawler): relative URL handling on non-start pages (#893) Gergő Móricz 2024-11-12 18:20:53 +0100
  • d430cfcbfb Update extract.ts Nicolas 2024-11-12 12:17:48 -0500
  • 5bbbb52a30 Update fireEngine.ts Nicolas 2024-11-12 12:17:03 -0500
  • 740a429790 feat(api): graceful shutdown for less 502 errors Gergő Móricz 2024-11-12 18:10:24 +0100
  • 540d4e5b56 Nick: Nicolas 2024-11-12 12:10:18 -0500
  • c327d688a6 fix(queue-worker): don't log timeouts Gergő Móricz 2024-11-12 18:10:11 +0100
  • 9f8b8c190f feat(scrapeURL): log URL for easy searching Gergő Móricz 2024-11-12 17:54:48 +0100
  • e95b6656fa fix(scrapeURL): don't log fetch request Gergő Móricz 2024-11-12 17:53:44 +0100
  • f42740a109 fix(scrapeURL): don't log engineResult Gergő Móricz 2024-11-12 17:52:32 +0100
  • 1ddace3a0f fix(crawl): further fixing mog/fix-relative-url-crawl Móricz Gergő 2024-11-12 16:39:08 +0100
  • 3815d24628 fix(scrape): better timeout handling Móricz Gergő 2024-11-12 13:16:40 +0100
  • aa9a47bce7 fix(queue-worker): logging job on batch scrape error Móricz Gergő 2024-11-12 13:00:19 +0100
  • 91f52287db feat(batchScrape): handle timeout Móricz Gergő 2024-11-12 12:42:39 +0100
  • f6db9f1428 fix(crawl-redis): batch scrape lockURL Móricz Gergő 2024-11-12 11:52:34 +0100
  • f2eb3a2d9a fix(crawler): relative URL handling on non-start pages Gergő Móricz 2024-11-11 22:18:30 +0100
  • d8bb1f68c6 fix(tests): maxDepth tests Gergő Móricz 2024-11-11 22:10:19 +0100
  • 68c9615f2d fix(crawl/maxDepth): fix maxDepth behaviour Gergő Móricz 2024-11-11 22:02:17 +0100
  • 7d576d13bf Update package.json Nicolas 2024-11-11 15:42:10 -0500
  • a8dc75f762
    feat(crawl): add parameter to treat differing query parameters as different URLs (#892) Gergő Móricz 2024-11-11 21:36:22 +0100
  • 4df7390e8c add code to make it work Gergő Móricz 2024-11-11 21:35:05 +0100
  • 5cb46dc494 fix(html-to-markdown): build error Gergő Móricz 2024-11-11 21:09:27 +0100
  • 2ca22659d3 fix(scrapeURL/llmExtract): fix schema-less LLM extract Gergő Móricz 2024-11-11 21:07:37 +0100
  • 56bebc8107 fix(html-to-markdown): reduce logging frequency Gergő Móricz 2024-11-11 20:53:21 +0100
  • d13a2e7d26 fix(scrapeURL): reduce logs Gergő Móricz 2024-11-11 20:51:29 +0100
  • 219f4732a0
    Merge pull request #881 from mendableai/fix/scroll-action Nicolas 2024-11-11 14:50:08 -0500
  • ddbf3e45a3
    Update package.json fix/scroll-action Nicolas 2024-11-11 14:49:50 -0500
  • 766377621e
    Merge pull request #880 from mendableai/python-sdk/next-handler Nicolas 2024-11-11 14:48:30 -0500
  • 9688bad60d
    Update __init__.py python-sdk/next-handler Nicolas 2024-11-11 14:48:11 -0500
  • 56a1ac07a4
    Merge pull request #878 from mendableai/mog/deduplicate-urls Nicolas 2024-11-11 14:33:13 -0500
  • 8e4e49e471 feat(generateURLPermutations): add tests Gergő Móricz 2024-11-11 20:29:17 +0100
  • 027aab31d2 fix(sitemap): scrape with tlsclient Gergő Móricz 2024-11-11 19:44:32 +0100