fix(scrape-single-url): return the only dataset item after the run is finished - #3
Merged
oleksandravalko merged 7 commits intoAug 6, 2025
Conversation
drobnikj
requested changes
Jul 22, 2025
| name: "Get Dataset Items", | ||
| description: "Returns data stored in a dataset. [See the documentation](https://docs.apify.com/api/v2/dataset-items-get)", | ||
| version: "0.0.2", | ||
| version: "0.0.3", |
Member
There was a problem hiding this comment.
Do we really need to change version of all actions, when we are updating just scrape singe URL?
Author
There was a problem hiding this comment.
I didn't get to the bottom of it, but this is the step I needed to take for PR checks to pass. I doesn't make much sense to me either, maybe this problem is caused by the fact, that this a PR from a different repo.
…t of terminal statuses to stop the loop
drobnikj
approved these changes
Jul 30, 2025
drobnikj
left a comment
Member
There was a problem hiding this comment.
Please import consts from package otherwise fine, pre-approving. 👍
protoss70
approved these changes
Jul 30, 2025
protoss70
left a comment
There was a problem hiding this comment.
I just have one comment about the polling rates, other than that looks good
… in between calls
matyascimbulka
added a commit
that referenced
this pull request
Sep 24, 2025
* feat(apify): Show list of built tags in actor run action * feat(apify): Allow selecting actor search source for run Actor action and change how Actor or task name is displayed * feat(apif): Remove wait for finish prop in run Actor action * fix(apify): Fix PR issues * fix(apif): Fix PR issues * fix(scrape-single-url): return the only dataset item after the run is finished (#3) * fix(scrape-single-url): return the only dataset item after the run is finished * fix(scrape-single-url): version up * fix(scrape-single-url): version up * fix(scrape-single-url): version up * fix(scrape-single-url): introduce a job status constant, expand a list of terminal statuses to stop the loop * fix(scrape-single-url): import constants from package, decrease delay in between calls * Migrate to use Apify client (#6) * feat(apify): Replace Axios with Apify client * fix(general): adding custom headers to client()=> preserve whole config to be passed to Axios later * feat(general): add linter script * fix(apify-get-dataset-items): a function for getting items and parsing of a result * fix(apify-run-actor): working sync and async, dynamic input schema injection, KVS output retrieval tested only string * fix(general): change maxResults for limit as an input field * fix(run-task-sync): move items retrieval to the component, add waitSecs determined by input or plan to prevent blunt timeout error, have the item retrieval logic be connected to run status, clean return value * fix(apify-scrape-single-url): incorporate timeouts, rework the whole API interaction logic * fix(apify-set-key-value-store-record): detection of content type, fixed API interaction * fix(apify-scrape-single-url): remove waiting timeout, return only dataset item, remove extra input fields connected to WCC run * fix(apify-run-actor): success message * fix(apify-run-task-synchronously): remove waiting for run to finish timeout * fix(app): remove paidPlan input filed config --------- Co-authored-by: Matyas Cimbulka <matyas.cimbulka@apify.com> * fix(apify-get-dataset-items) 6: change input parameters (#8) * chore(apify): Bump component versions * chore: Sync upstream repo (#9) * feat(apify): Prefill values from the input schema (#7) * Revert "chore: Sync upstream repo (#9)" This reverts commit cd804ba. * Revert "chore(apify): Bump component versions" This reverts commit 6040822 which for some reason bumped version of the wrong components. * fix(apify): Fix build tag * fix(apify): Address issues in run task synchronously action * feat(apify): Add default crawler type to scrape single url * fix(apify): Address issues from PR * chore(apify): Change component versions * fix(apify): Fix run Actor action * fix(apify): Fix run get dataset items action * fix(apify): Fix typos for PR --------- Co-authored-by: Oleksandra Valko <oleksandra.valko@apify.com>
drobnikj
pushed a commit
that referenced
this pull request
Sep 3, 2026
* feat(super_carl) add new components Summary Fixes review comments from PipedreamHQ#21252 since we are unable to make the changes directly in their fork as maintainerContribution is disabled by the submitted user. * Changes with respect to description and coderabbit review comment Changes with respect to description and coderabbit review comment * changes with respect to description and utils file changes with respect to description and utils file * Changes with respect to review comment Changes with respect to review comment * Add field selection to search actions, auto-generate draft session id - search-people/search-companies/search-jobs/search-posts: add an additive `fields` prop to return only named fields per row, since full rows carry deep connection/evidence metadata large enough to exceed the MCP output ceiling on a single call. - create-communication-draft: make agentSessionId optional and auto-generate one when omitted, since callers have no way to know what a valid session id looks like. - search-people/search-jobs: description guidance clarifying the people database is read-only (no delete/edit) and that job location filters scope the hiring company/people, not the job posting. * Add field selection to check-communication-capabilities Each channel entry carries verbose reason/relationship metadata that pushed this tool's output to ~9k chars average, flagged by the eval suite's oversized-output check. Add the same additive `fields` prop already used by the search actions, scoped to the `channels` array. * Changes with respect to coderabbit review Changes with respect to coderabbit review * Removed communication-draft component and added ai-optimised Removed communication-draft component and added ai-optimised * Update readme file Update readme file * Strip search-people's internal debug telemetry from every response Evals #1/#9 in evals/super_carl kept hitting the 25k-token MCP ceiling even with `fields` + `limit:1`. Traced it to `search_metadata` (45-55k chars) and `entity_resolution` (~10k chars) — Elasticsearch execution internals (query embeddings, bitmap counts, raw es_response, filter provenance) that ship at the top level of every people-search response regardless of `fields`, which only scopes each row in `users`. Neither field is documented or actionable for a caller, so both are dropped unconditionally rather than gated behind an opt-in. * Fix search-posts people bloat, correct wrong Fields examples - search-posts: drop matched_posts from each deduped person row unconditionally (~60% of that row's size, measured 1004/1687 bytes on a real response, and redundant with `results` which already has the matching posts) - `fields` alone couldn't fix this since it only scoped `results`, and people/post field vocabularies don't overlap anyway. Contributed to evals #4/#9 hitting the MCP output ceiling even with `fields` set on the posts side. - search-posts, search-jobs: description examples named nested paths (`author.name`, `company.name`) that don't exist - the real fields are flat (`author_name`, `company_name`). Confirmed from real responses: the model tried the documented dot-path first, got a silently empty field, and burned a retry figuring out the flat name each time (evals #3, #4). - search-people: evals #1 and #9 each hit the ceiling on a first attempt that combined Preview off + `relationshipDetail: intro_paths` + `evidenceFormat` with no Fields, then recovered on retry once Fields was added. Description now says to pass Fields proactively with that combination instead of waiting for a truncated first call. * addressing coderaabit internal review comment addressing coderaabit internal review comment * Addressed review comments Addressed review comments * Updated description Updated description * fix(super_carl): stop search-people from blowing the MCP output ceiling search-people's raw API response carries a search_metadata field with internal search-engine telemetry (query execution plans, raw Elasticsearch response, bitmap counts) that measured 46KB on a single-row response, independent of limit/fields/preview. Strip it the same way entity_resolution already is. Also add a safe default field set for preview:false combined with relationshipDetail/evidenceFormat, which the tool's own description already warns can exceed the size limit even for one row (measured ~15KB/row). This only applies when the caller passed no explicit fields, so it's a fallback, not a behavior change for anyone already using fields. Verified via eval re-runs: fixed the two evals that were hitting the ceiling outright (network sync status check, engineering jobs search timing out from retries). The remaining two evals that exercise this path could not be re-verified after this last change - Super Carl's free-tier API credits (5/30-day cycle) were exhausted by the debugging session's own probing. * fix(super_carl): prevent search-people truncation and steer search-companies away from resolveOnly for size/description search-people's truncation safety net only applied when preview was explicitly false, missing the preview:true + relationshipDetail:summary combination that also exceeds the MCP output ceiling — broaden it to trigger on relationshipDetail alone. search-companies' resolveOnly mode returns identity-disambiguation data only; the model was using it for "size and what they do" questions and had no tool-grounded data to answer with. Description now says explicitly what resolveOnly does and doesn't return, and to use resultMode: detailed instead. Found via eval-driven testing (evals/super_carl/evals.yaml, PR PipedreamHQ#21630).
drobnikj
pushed a commit
that referenced
this pull request
Sep 3, 2026
…ages (PipedreamHQ#21684) * feat(slack_v2): read direct messages by user ID + deprecate List Messages Iterated against the MCP eval suite (evals/slack_v2); this ships the fixes the evals surfaced. Suite green on Sonnet 5 (3/3, pass^2) after the change. - slack_v2.app.mjs: eval PipedreamHQ#36 ("read my DMs with myself") failed — the read tools forwarded a `U…` user id straight to conversations.history, which only accepts a conversation id and answered channel_not_found (writing to a DM already worked via chat.postMessage auto-open). Added openConversation() (conversations.open) and made resolveChannelId open the DM for a user id — the read-side counterpart to posting — so history, thread-replies, and reactions now accept a user id. Also made the id regexes case-sensitive (Slack ids are uppercase-only) so an all-alphanumeric lowercase channel NAME isn't misclassified as an id. PipedreamHQ#36 FAIL→PASS. [shared by 10 actions] - get-channel-history: description + `channel` prop now document reading a DM by user id. [minor] - list-messages: the legacy twin of Get Channel History was winning routing on the channel-read evals (#3/#25 warned expected_tools_missing, precision 0%). Renamed to "List Messages (Deprecated)" and steered to Get Channel History (name + first line are the tool-search retrieval key); run() unchanged, existing workflows still work. #3/#25/PipedreamHQ#36 → pass^2 3/3. [patch] - get-thread-replies, browse-files, set-channel-topic, get-channel-details, invite-user-to-channel, delete-message, add-reaction, edit-message, list-members-in-channel: version-only bumps for the shared resolveChannelId change. [patch] App package.json bumped 0.7.0 → 0.8.0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(slack_v2): bump remaining component versions for shared app-file change The slack_v2.app.mjs change in this PR touches a shared dependency file, so CI's version check flags every component in the app. Patch-bump the remaining actions and sources (the resolveChannelId consumers were already bumped in the prior commit) to satisfy the check. No behavior change in these files. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * add ai-optimized marker --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
drobnikj
pushed a commit
that referenced
this pull request
Sep 3, 2026
…1826) * feat(dappier): AI-optimized Dappier action set for MCP (real-time search, recommendations, analytics) Initial AI-optimized Dappier components for the MCP tool surface, covering issue PipedreamHQ#21660 (real-time web search, AI content recommendations, data-model querying) plus the four Analytics API endpoints. Iterated against the MCP eval suite (evals/dappier) — 9/10 green on Sonnet 5 (trials: 1); the one failure is an upstream HTTP 500 on POST /app/v2/search, not a component defect. agent-audit: 100/100. - search-real-time-data (new, 0.0.1): real-time web/data search via a Dappier AI model (am_ id); returns a synthesized answer. Eval #9 passes. - get-ai-recommendations (new, 0.0.1): AI-ranked content recommendations for a data model (dm_ id); optional additive `fields` projection trims large article payloads. Eval #10 blocked by an upstream 500 (server-side). - get-ask-ai-analytics (new, 0.0.1): aggregate Ask AI widget analytics. Evals #1/#4/#5 pass. - get-ask-ai-logs (new, 0.0.1): raw Ask AI conversation logs with page/limit pagination + paging guidance. Evals #2/#6 pass. - get-sponsored-conversations-analytics (new, 0.0.1): sponsored-conversation (ad campaign) analytics. Eval #7 passes. - get-session-intelligence (new, 0.0.1): session intent/topic breakdowns. Evals #3/#8 pass. - dappier.app.mjs: shared analytics prop definitions + GET request methods. - common/utils.mjs: validateDateRange (365-day cap) + pluckFields projection helper. - common/constants.mjs: analytics interaction types, range + page-size bounds. App package.json bumped 0.0.1 -> 0.1.0 (minor -- new actions). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(dappier): add trailing newline to package.json (eol-last) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(dappier): address review — numArticlesRef max, own-prop projection, date default wording - get-ai-recommendations: numArticlesRef description said max 1000 but schema is max 100; corrected the doc. - common/utils.mjs: pluckFields now uses Object.hasOwn so a requested field name that collides with an inherited prop (e.g. toString) is not copied; own result fields still are. - dappier.app.mjs: reworded startDate default from the brittle/off-by-one '7 days before today' to 'the last 7 days (UTC), i.e. today and the six prior days', matching observed API behavior. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(dappier): resolve one-sided analytics date windows before sending Verified against the live API: omitting either start_date or end_date makes Dappier reset BOTH bounds to its default trailing-7-day window, silently discarding the bound the caller supplied — so a one-sided range returned the wrong period. resolveDateRange() now fills the missing bound (missing end -> today UTC; missing start -> 6 days before the end) and always sends both, so a supplied bound is honored; both-omitted still defers to the API default. All four analytics actions use it; start/end prop docs updated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(dappier): match analytics date-range cap to the API (inclusive days) Probed the live API: it accepts a 364-day start/end difference (365 inclusive days) and returns 400 at a 365-day difference (366 inclusive days). validateDateRange used '> 365 days difference', so a 365-day-difference window slipped past the fail-fast and hit a raw API 400. Now counts inclusive days (difference + 1) and caps at 365, matching the API exactly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(dappier): note session_intelligence wrapper in get-session-intelligence description The API nests all six breakdowns under a single top-level `session_intelligence` object (verified live). The description listed them as if they were root keys, so an agent would look for them at the root and miss the `session_intelligence.` prefix. Description-only; no behavior change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(dappier): use camelCase widgetId in session-intelligence example The example told the agent to pass `widget_id`, but the input prop is `widgetId` (run() maps it to the `widget_id` query param). Match the example to the prop the agent actually sets. Description-only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(dappier): use camelCase prop names in remaining action examples Same fix as get-session-intelligence, applied to the sibling descriptions: the agent-facing 'Example:' hints used API param names (start_date, end_date, campaign_id, data_model_id) instead of the camelCase input props (startDate, endDate, campaignId, dataModelId). API-mapping mentions stay snake_case. Description-only; no behavior change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(dappier): camelCase input-guidance for model-id props The 'Provide a ...' input guidance used API param names — get-ai-recommendations said 'Provide a data_model_id' (prop is dataModelId) and search-real-time-data said 'Provide an ai_model_id' (prop is aiModelId). Use the camelCase input keys; snake_case is retained only where describing the API request mapping (query param / path template). Description-only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Adding missing dependencies field --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: GTFalcao <gtfalcao96@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
WHY
Connected to this issue: https://github.com/orgs/apify/projects/19/views/1?pane=issue&itemId=118956431&issue=apify%7Cintegrations-team%7C4