Eric Ciarla
2b40729cc2
Update index.test.ts
2024-06-15 08:56:32 -04:00
Eric Ciarla
f22759b2e7
Update index.test.ts
2024-06-14 19:42:11 -04:00
Eric Ciarla
a6b7197737
Fix for maxDepth
2024-06-14 19:40:37 -04:00
Nicolas
4ec863718b
Merge pull request #283 from mendableai/nsc/crawler-fixes
...
Fixes crawler getting confused with base paths that contain www.
2024-06-14 13:50:32 -07:00
Nicolas
43767360d8
Merge branch 'main' into nsc/rate-limiter-tests
2024-06-14 13:50:21 -07:00
Nicolas
e88cb314c8
Update crawler.ts
2024-06-14 13:44:54 -07:00
Rafael Miller
361cba4119
Merge pull request #175 from mendableai/test/load-testing
...
Test/load testing
2024-06-14 17:39:01 -03:00
Nicolas
7b11ace87d
Create rate-limiter.test.ts
2024-06-14 12:31:42 -07:00
rafaelsideguide
e369d1dd0e
Update index.test.ts
2024-06-14 16:17:54 -03:00
Nicolas
e37aa3db57
Nick: fixed rate limit on status
2024-06-14 12:13:02 -07:00
rafaelsideguide
a6ed2e693f
Update index.test.ts
2024-06-14 15:22:52 -03:00
rafaelsideguide
ad7795f973
Merge remote-tracking branch 'origin/main' into test/load-testing
2024-06-14 15:14:01 -03:00
rafaelsideguide
354712a8a3
just changed the name for the test?
2024-06-14 13:02:04 -03:00
Eric Ciarla
2c5f5c0ea2
Merge branch 'main' into feat/maxDepthRelative
2024-06-14 11:49:12 -04:00
Eric Ciarla
80c10393b4
Update index.test.ts
2024-06-14 11:32:30 -04:00
Eric Ciarla
42ed1f4479
Update index.test.ts
2024-06-14 11:20:24 -04:00
Eric Ciarla
8830acce07
Update index.test.ts
2024-06-14 11:11:58 -04:00
Eric Ciarla
278bb311cb
Update index.test.ts
2024-06-14 11:02:39 -04:00
Eric Ciarla
36a62727b8
Update index.test.ts
2024-06-14 10:52:43 -04:00
Rafael Miller
f9c7ca9388
Merge branch 'main' into feat/issue-266
2024-06-14 11:47:58 -03:00
Rafael Miller
3e2e76311c
Merge branch 'main' into feat/issue-205
2024-06-14 11:25:20 -03:00
Eric Ciarla
59451754f5
Add tests
2024-06-14 10:14:07 -04:00
Eric Ciarla
9b254c1cd0
Update index.test.ts
2024-06-14 09:48:14 -04:00
Eric Ciarla
9aba451b18
Update index.test.ts
2024-06-14 09:33:43 -04:00
rafaelsideguide
5dd18ca79b
fixed edge cases
2024-06-14 09:46:55 -03:00
Eric Ciarla
ab9de0f5ab
Update maxDepth tests
2024-06-13 18:46:30 -04:00
Eric Ciarla
393bd45237
Update index.test.ts
2024-06-13 18:13:15 -04:00
Eric Ciarla
71c98d8b80
Update logic
2024-06-13 18:00:52 -04:00
Eric Ciarla
095951aa4d
Update test
2024-06-13 17:40:00 -04:00
Eric Ciarla
5e8aa92788
Update index.ts
2024-06-13 17:33:13 -04:00
Eric Ciarla
bf10e9d392
Update index.test.ts
2024-06-13 17:28:59 -04:00
Eric Ciarla
65d63bae45
Update index.ts
2024-06-13 17:17:44 -04:00
Eric Ciarla
32e814bedc
Update index.ts
2024-06-13 17:02:30 -04:00
Nicolas
6fc1ee32fd
Merge pull request #275 from mendableai/feat/issue-273
...
Added pageOptions.removeTags
2024-06-13 13:27:01 -07:00
rafaelsideguide
bb859ae9a7
Added metadata.pageStatusCode and metadata.pageError properties to the responses
2024-06-13 17:08:40 -03:00
rafaelsideguide
676d6e8ab5
Added pageOptions.removeTags
2024-06-13 10:51:05 -03:00
Nicolas
182f8d4d6c
Update index.ts
2024-06-12 18:07:05 -07:00
Nicolas
11b6d5afa5
Update fly.toml
2024-06-12 18:00:22 -07:00
Nicolas
67dc46b454
Nick: clusters
2024-06-12 17:53:04 -07:00
rafaelsideguide
d20af257ba
Added jobId to webhook data
2024-06-12 15:38:41 -03:00
rafaelsideguide
e37d151404
added parsePDF option to pageOptions
...
user can decide if they are going to let us take care of the parse or they are going to parse the pdf by themselves
2024-06-12 15:06:47 -03:00
rafaelsideguide
01c9f071fa
fixed
2024-06-12 11:27:06 -03:00
rafaelsideguide
dc6acbf1f0
Merge remote-tracking branch 'origin/main' into feat/allowbackwardcrawling-option
2024-06-12 11:01:05 -03:00
Nicolas
f93231499f
Merge pull request #265 from mendableai/feat/issue-264
...
[Feat] Added route to clean completed jobs and a github action cron that triggers every 24h
2024-06-11 21:33:52 -07:00
Nicolas
45dee63943
Merge pull request #262 from mendableai/nsc/webhook-self-host-fix
...
Only fetch webhook from db if self host webhook not set and using db auth
2024-06-11 15:46:57 -07:00
rafaelsideguide
157fbe4a1e
added bull auth key
2024-06-11 17:52:01 -03:00
rafaelsideguide
df3a678cf4
getting back the cancel test, this should work
2024-06-11 17:46:56 -03:00
rafaelsideguide
def2ba9987
added tests
2024-06-11 17:46:25 -03:00
Nicolas
1e3e06a1d5
Update replacePaths.test.ts
2024-06-11 13:02:39 -07:00
Nicolas
2239e03269
Update replacePaths.test.ts
2024-06-11 12:54:02 -07:00
Nicolas
520739c9f4
Nick: fixed bugs associated with absolute path replacements
2024-06-11 12:43:16 -07:00
Nicolas
b87725c683
Update openapi.json
2024-06-11 12:08:49 -07:00
rafaelsideguide
ee282c3d55
Added allowBackwardCrawling option
2024-06-11 15:24:39 -03:00
rafaelsideguide
a9f93c2f1e
Added route to clean completed jobs and a github action cron that triggers every 24h
2024-06-11 14:18:05 -03:00
Nicolas
da38dad9a7
Merge branch 'main' of https://github.com/mendableai/firecrawl
2024-06-10 18:26:31 -07:00
Nicolas
9390816c1b
Update openapi.json
2024-06-10 18:26:25 -07:00
Nicolas
f6b06ac27a
Nick: ignoreSitemap, better crawling algo
2024-06-10 18:12:41 -07:00
Nicolas
1bd0327e1a
Merge branch 'main' into nsc/pageoptions-crawler
2024-06-10 17:15:10 -07:00
Nicolas
99f2ffd6d5
Update webhook.ts
2024-06-10 17:03:10 -07:00
Nicolas
7ae9778642
Update single_url.ts
2024-06-10 16:57:31 -07:00
Nicolas
913c1dd568
Nick: fetch -> axios and fix timeouts
2024-06-10 16:49:03 -07:00
Nicolas
3091f0134c
Nick:
2024-06-10 16:27:10 -07:00
Nicolas
f24ca76618
Nick: removing rate limit emails for now
2024-06-07 10:39:11 -07:00
Nicolas
98d82c4cec
Update search.ts
2024-06-06 20:02:21 -07:00
Nicolas
5e80f8af87
Nick: llm extract 50
2024-06-06 18:35:44 -07:00
rafaelsideguide
f5318ea7d7
Update index.test.ts
2024-06-06 16:50:20 -03:00
rafaelsideguide
cd7f9abcec
Update index.test.ts
2024-06-06 16:44:46 -03:00
rafaelsideguide
7b9b668b95
Update index.test.ts
2024-06-06 16:36:51 -03:00
rafaelsideguide
82e0ed4cd3
Update index.test.ts
2024-06-06 16:33:27 -03:00
rafaelsideguide
dac7612be2
Merge branch 'main' of https://github.com/mendableai/firecrawl into 194-sdk-ci-pipeline-for-publishing-pythonnode-sdk
2024-06-06 16:07:25 -03:00
Nicolas
c2ad358390
Nick:
2024-06-06 12:05:20 -07:00
rafaelsideguide
79ec9f04dc
Merge branch 'main' of https://github.com/mendableai/firecrawl into 194-sdk-ci-pipeline-for-publishing-pythonnode-sdk
2024-06-06 15:58:14 -03:00
Nicolas
de06b13deb
Update rate-limiter.ts
2024-06-06 11:56:22 -07:00
Nicolas
27a8fd0c3c
Update rate-limiter.ts
2024-06-06 11:56:00 -07:00
Nicolas
1129d33321
Update rate-limiter.ts
2024-06-06 11:53:12 -07:00
rafaelsideguide
b234b4be5a
Merge branch 'main' into 194-sdk-ci-pipeline-for-publishing-pythonnode-sdk
2024-06-06 15:44:29 -03:00
rafaelsideguide
af0bfca847
Merge branch 'main' into 194-sdk-ci-pipeline-for-publishing-pythonnode-sdk
2024-06-06 15:36:28 -03:00
rafaelsideguide
8132f22c73
nice
2024-06-06 15:36:20 -03:00
Nicolas
f1b5ec8517
Nick: fixes
2024-06-06 11:23:10 -07:00
Nicolas
deae7dcd61
Update email_notification.ts
2024-06-06 10:41:54 -07:00
Nicolas
f725fa5a97
Update email_notification.ts
2024-06-06 10:41:23 -07:00
Nicolas
0310da6729
Update rate-limiter.ts
2024-06-06 09:31:44 -07:00
Nicolas
01503c1fbf
Nick:
2024-06-06 09:29:25 -07:00
Nicolas
525b4f2a83
Update rate-limiter.ts
2024-06-05 14:38:10 -07:00
Nicolas
d7f8208cdb
Update email_notification.ts
2024-06-05 13:53:31 -07:00
Nicolas
ec10eb09f3
Update credit_billing.ts
2024-06-05 13:22:03 -07:00
Nicolas
5991000d2b
Update credit_billing.ts
2024-06-05 13:21:15 -07:00
Nicolas
5683bb2cc8
Nick:
2024-06-05 13:20:26 -07:00
rafaelsideguide
164676c70a
bugfix screenshot for readme pages
2024-06-05 15:34:42 -03:00
Nicolas
b4c6819a54
Nick:
2024-06-05 11:11:09 -07:00
rafaelsideguide
0d51b11dcd
missing breaks
2024-06-05 15:02:28 -03:00
Nicolas
beb7526d1d
Update webhook.ts
2024-06-05 10:38:05 -07:00
Nicolas
1a16378fe8
Merge pull request #234 from JakobStadlhuber/feat/webhook-self-hosted
...
Add support for Self-Hosted Webhook URL Usage and added project_id into the webhook payload
2024-06-05 10:25:05 -07:00
Nicolas
7cb14edec8
Nick:
2024-06-05 10:13:52 -07:00
Rafael Miller
9e000ded03
Merge branch 'main' into feat/better-gdrive-pdf-fetch
2024-06-05 14:07:56 -03:00
rafaelsideguide
ccc55127d6
Added scroll xpaths on fire-engine for handling readme docs
2024-06-05 11:48:41 -03:00
rafaelsideguide
b5045d1661
[feat] improved the scrape for gdrive pdfs
2024-06-04 17:47:28 -03:00
Nicolas
96257b7b17
Update handleCustomScraping.ts
2024-06-04 12:22:46 -07:00
Nicolas
674500affa
Nick:
2024-06-04 12:15:39 -07:00
rafaelsideguide
5ae4d1caf5
Update single_url.ts
2024-06-04 15:28:09 -03:00
Jakob Stadlhuber
9e5ddec207
Remove default webhook URL from .env.example
...
The default value for the SELF_HOSTED_WEBHOOK_URL in the .env.example file was removed to prevent unintentional exposure or usage. The users are now required to explicitly specify
2024-06-04 19:56:35 +02:00
Jakob Stadlhuber
6208f4207d
Add support for Self-Hosted Webhook URL Usage and added project_id into the webhook payload
...
This commit introduces the capability of using a Self-Hosted Webhook URL. The application now checks for a self-hosted URL before querying the database for the webhook settings. If a Self-Hosted Webhook URL is set in the environment variables, it will be used directly, diminishing unnecessary database queries.
2024-06-04 19:55:07 +02:00
rafaelsideguide
64a4338ff0
Update single_url.ts
2024-06-04 14:40:05 -03:00
Rafael Miller
02fe470e20
Merge pull request #148 from mendableai/nsc/improvemnts-fixes-misc
...
Better fallbacks for initial crawl start
2024-06-04 14:31:10 -03:00
Rafael Miller
b80fb374e5
Merge branch 'main' into playwright-service-bug-222
2024-06-04 11:57:17 -03:00
rafaelsideguide
6920ec8a61
bugfixing. already on main
2024-06-04 11:05:50 -03:00
Nicolas
d91b725c6f
Update fly.toml
2024-06-04 00:41:15 -07:00
Nicolas
cbf8d79cce
Update pdfProcessor.ts
2024-06-04 00:13:37 -07:00
Nicolas
3fc9004ba8
Update fly.toml
2024-06-03 23:49:46 -07:00
Nicolas
2ea01f1456
Update single_url.ts
2024-06-03 23:42:39 -07:00
Nicolas
854d5b3cb3
Update single_url.ts
2024-06-03 23:32:55 -07:00
Nicolas
99059814a8
Nick:
2024-06-03 21:32:48 -07:00
Nicolas
918059ee9e
Merge branch 'main' into nsc/improvemnts-fixes-misc
2024-06-03 16:46:02 -07:00
Nicolas
38e583f66c
Update socialBlockList.test.ts
2024-06-03 16:44:23 -07:00
Nicolas
c69c89f838
Nick:
2024-06-03 16:42:42 -07:00
Nicolas
48d1ec05b2
Merge branch 'main' into nsc/improved-blocklist
2024-06-03 16:38:03 -07:00
Nicolas
d30ced4394
Merge pull request #221 from mendableai/nsc/fwd-header-auth
...
feat: Ability to forward headers to reliable providers for auth etc...
2024-06-03 16:33:40 -07:00
Romain Bruyère
4987f901d1
Merge branch 'mendableai:main' into main
2024-06-03 21:29:33 +02:00
rombru
3ff91ddd1f
fix: use @ instead of # for default BULL_AUTH_KEY. hash mark is reserved for URI fragments.
2024-06-03 21:28:25 +02:00
rafaelsideguide
c1aed1360e
Update index.test.ts
2024-06-03 15:51:07 -03:00
rafaelsideguide
1fc3a15149
Update single_url.ts
2024-06-03 15:24:40 -03:00
Nicolas
fde522c3e1
Update single_url.ts
2024-06-02 20:23:45 -07:00
Matt Joyce
deefe65cbe
Change the way the playwright response is parsed
...
Was failing with a Type Error, but actually looked ok.
This fixes the type error, and stop scraper fallback.
2024-06-01 19:16:56 +10:00
Matt Joyce
14896a9fdd
Fix PLAYWRIGHT_MICROSERVICE_URL
...
It needs to end in html, otherwise scrape will 404
2024-06-01 19:03:16 +10:00
Nicolas
8cb62dde92
Update website_params.ts
2024-05-31 16:09:39 -07:00
Nicolas
3b8059edb6
Update single_url.ts
2024-05-31 15:43:06 -07:00
Nicolas
6bea803120
Nick:
2024-05-31 15:39:54 -07:00
Nicolas
2139129296
Nick: v12
2024-05-31 11:39:55 -07:00
Nicolas
260e31c68b
Merge branch 'nsc/new-pricing'
2024-05-30 16:08:31 -07:00
Nicolas
aa8133ca7f
Update load-testing-example.ts
2024-05-30 16:07:14 -07:00
Nicolas
0c115c6181
Merge pull request #216 from mendableai/nsc/new-pricing
...
feat: New pricing/limits changes
2024-05-30 15:36:59 -07:00
Nicolas
6860ace4af
Nick:
2024-05-30 15:07:49 -07:00
Nicolas
6ceb7ff50a
Nick:
2024-05-30 14:46:55 -07:00
Nicolas
33f10a7f91
Nick: fixes
2024-05-30 14:42:32 -07:00
Nicolas
ace46f340b
Nick: new limits, new pricing
2024-05-30 14:31:36 -07:00
Nicolas
6c939d534d
Nick: small refactor
2024-05-29 19:43:51 -07:00
Eric Ciarla
37915e11e8
Final push
2024-05-29 21:18:24 -04:00
Eric Ciarla
a0e404f94e
init commit
2024-05-29 18:56:57 -04:00
rafaelsideguide
ee9a2184e2
Added custom scraping conditions for readme docs
2024-05-29 13:39:43 -03:00
Nicolas
c20c38721d
Update index.test.ts
2024-05-28 17:17:20 -07:00
Nicolas
0f43a12906
Update index.test.ts
2024-05-28 17:17:12 -07:00
Nicolas
1b3547dcf2
Nick:
2024-05-28 12:56:24 -07:00
Nicolas
1ef307cb6f
Nick: checks
2024-05-27 10:01:12 -07:00
Nicolas
01cc91c53d
Update fly.staging.toml
2024-05-27 10:00:52 -07:00
Nicolas
1bbfb98d7e
Merge pull request #186 from Keredu/main
...
Limit on /search is not deterministic
2024-05-26 18:08:16 -07:00
Nicolas
7e2df7bd5e
Update auth.ts
2024-05-26 18:07:21 -07:00
Simon H
115204e6b6
Feat: Provide more details for 429 error msg
...
- Added better error code for when rate limit exceeded including
consumed/remaining points, reset date and retry-after seconds
2024-05-25 12:03:20 -04:00
Keredu
2192978f91
Limit on /search is not deterministic
2024-05-25 00:12:26 +02:00
Nicolas
e98434606d
Update blocklist.ts
2024-05-24 15:04:15 -07:00
Nicolas
e5c8719554
Update blocklist.ts
2024-05-24 14:53:04 -07:00
rafaelsideguide
d39860c08b
Merge branch 'main' into feat/idempotency-key
2024-05-24 14:15:37 -03:00
Nicolas
53a214cefb
Merge pull request #168 from mendableai/nsc/allowed-keywords-in-blocklist
...
feat: Allow privacy/legal/ other pages in social media websites
2024-05-24 09:43:15 -07:00
Jakob Stadlhuber
9fc5a0ff98
Update comment in .env.example for proxy settings
...
This commit modifies the comment in .env.example to specify that proxy settings are for Playwright. This clarification aims to provide users a more clear context about when and why these proxy settings are used.
2024-05-24 17:45:59 +02:00
Jakob Stadlhuber
b001aded46
Add proxy and media blocking configurations
...
Updated environment variables and application settings to include proxy configurations and media blocking option. The proxy settings allow users to use a proxy service, while the media blocking is an optional feature that can help save bandwidth. Changes have been made in the .env.example, docker-compose.yaml, and main.py files.
2024-05-24 17:41:34 +02:00
rafaelsideguide
35927a65a5
Merge branch 'main' into feat/idempotency-key
2024-05-23 12:20:06 -03:00
rafaelsideguide
184e4678f1
bugfix on idempotency key check
2024-05-23 11:47:04 -03:00
rafaelsideguide
aa6df4305e
crawl load tests 6 and 7
2024-05-22 18:20:24 -03:00
rafaelsideguide
73f1d09d39
Update website_params.ts
2024-05-22 15:07:12 -03:00
rafaelsideguide
4dfc371241
Update index.test.ts
2024-05-22 14:38:41 -03:00
rafaelsideguide
f4a3469b9e
Merge branch 'main' into bug/crawl-limit
2024-05-22 14:27:28 -03:00
Nicolas
0d187f0425
Merge pull request #77 from tractorjuice/patch-1
...
Add additional file extensions to crawler.ts
2024-05-22 10:16:49 -07:00
rafaelsideguide
04a0bef0fb
Merge branch 'main' into test/load-testing
2024-05-22 11:26:19 -03:00
rafaelsideguide
e4573c08ca
Update website_params.ts
2024-05-22 11:24:48 -03:00
rafaelsideguide
068a240ab4
load tests for scrape route
2024-05-22 09:30:32 -03:00
Nicolas
cb2bd0e71f
Update index.test.ts
2024-05-21 19:03:32 -07:00
Nicolas
253abb849f
Update rate-limiter.ts
2024-05-21 18:53:58 -07:00
Nicolas
229b9908d2
Nick: only enable hyper dx in prod
2024-05-21 18:52:46 -07:00
Nicolas
a8ff295977
Update single_url.ts
2024-05-21 18:50:42 -07:00
Nicolas
a5e718b084
Nick: improvements
2024-05-21 18:34:23 -07:00
Nicolas
6285f12cd1
Merge pull request #167 from mendableai/nsc/hyper-dx-integration
...
feat: HyperDX Integration
2024-05-21 13:19:38 -07:00
rafaelsideguide
75f4e34d8e
Merge branch 'main' into test/load-testing
2024-05-21 10:28:02 -03:00
rafaelsideguide
ec46065066
Update rate-limiter.ts
2024-05-21 10:07:27 -03:00
Nicolas
7f64fe884a
Update blocklist.ts
2024-05-20 17:26:01 -07:00
Nicolas
756f54466d
Nick: allowed keywords for now
2024-05-20 17:24:21 -07:00
Nicolas
01783dc336
Update openapi.json
2024-05-20 17:10:55 -07:00
Nicolas
77a79b5a79
Nick: max num tokens for llm extract (for now) + slice the max
2024-05-20 17:07:38 -07:00
Nicolas
2644e1c029
Update .env.example
2024-05-20 13:36:51 -07:00
Nicolas
9e61d431f0
Nick: hyper dx integration init
2024-05-20 13:36:34 -07:00
Nicolas
c74f757b53
Update rate-limiter.ts
2024-05-19 13:05:36 -07:00
Nicolas
98a39b39ab
Nick: increased rate limits
2024-05-19 12:59:29 -07:00
Nicolas
18fa15df25
Update index.test.ts
2024-05-19 12:50:06 -07:00
Nicolas
614c073af0
Nick: improvements
2024-05-19 12:45:46 -07:00
Nicolas
f473793ba3
Merge branch 'main' into feat/rate-limits
2024-05-19 12:23:34 -07:00
Nicolas
4efebf7a4b
Merge branch 'test/load-testing' of https://github.com/mendableai/firecrawl into test/load-testing
2024-05-19 12:22:51 -07:00
Nicolas
5792cd022c
Update fly.staging.toml
2024-05-19 12:22:49 -07:00
rafaelsideguide
d667e1417b
added fly staging load test
...
- being rate limited. Need to add the token to the rate-limit functions
2024-05-17 19:09:19 -03:00
Nicolas
7630565c26
Create fly.staging.toml
2024-05-17 14:33:59 -07:00
rafaelsideguide
a480595aa7
Update index.test.ts
2024-05-17 15:41:27 -03:00
rafaelsideguide
54049be539
Added e2e tests
2024-05-17 15:37:47 -03:00
Nicolas
6feb21cc35
Update website_params.ts
2024-05-17 11:21:26 -07:00
Nicolas
5be208f595
Nick: fixed
2024-05-17 10:40:44 -07:00
Nicolas
eb88447e8b
Update index.test.ts
2024-05-17 10:00:05 -07:00
Nicolas
df6c3d1e7d
Merge branch 'main' into detect-pdfs
2024-05-17 09:55:51 -07:00
Nicolas
9d635cb2a3
Nick: docx support
2024-05-16 11:48:02 -07:00
Nicolas
bcce0544e7
Update openapi.json
2024-05-16 11:03:32 -07:00
Nicolas
80250fb54f
Update index.test.ts
2024-05-15 17:40:46 -07:00
Nicolas
098db17913
Update index.ts
2024-05-15 17:37:09 -07:00
Nicolas
93b1f0334e
Update index.test.ts
2024-05-15 17:35:06 -07:00
Nicolas
123fb784ca
Update index.test.ts
2024-05-15 17:29:22 -07:00
Nicolas
4a6cfb6097
Update index.test.ts
2024-05-15 17:22:29 -07:00